PaddleOCR

开源 开源免费 搜索知识
GitHub 星标 ★ 87296
维护状态 较活跃
是否开源
定价模式 开源免费

项目数据

分类搜索知识
开发团队PaddlePaddle
所属国家
定价模式开源免费
价格区间
是否开源
开源协议Apache-2.0
主要语言Python
技术栈/模型ai4science,chineseocr,document-parsing,document-translation,kie,ocr,paddleocr-vl,pdf-extractor-rag,pdf-parser,pdf2markdown,pp-ocr,pp-structure,rag
GitHub 星标★ 87296
30天Star增速
HF 下载量
上线时间2020-05-08 00:00:00
最近更新2026-08-09 00:00:00
维护状态较活跃
中文支持
访问方式
移动端支持
综合评分
收录时间2026-08-09
浏览次数0

核心亮点

    不足之处

      适用场景

        替代项目

        项目介绍

        PaddleOCR 是一个搜索知识领域的开源项目,官方简介:Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.。项目使用 Python 开发,在 GitHub 上获得 87285 星标。

        上一篇:awesome-llm-apps

        下一篇:hello-agents

        同类项目推荐

        firecrawl 开源

        网页抓取像喝水一样简单,开发者省下整周加班

        The context API to search, scrape, and interact with the web at scale.

        ★ 164060 2026-08-09
        VibeSearchBench 开源

        The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-h···

        ★ 409 2026-08-09
        OpenSearch-VL 开源

        OpenSearch-VL provides a fully open recipe for training strong multimodal deep searc···

        ★ 262 2026-08-09