← プロジェクト一覧に戻る

PaddleOCR

スキャン画像や写真、PDFを、100言語以上に対応しながらCPUでも動く形でMarkdownやJSONに変換する

Apache-2.0
スター
88.3k
フォーク
11.2k
オープンIssue
234
最終コミット
2026年7月22日

PaddleOCRとは

自社の文書を対象にした検索は、モデルにたどり着く前に失敗していることがほとんどです。表のスキャンや2段組のPDFが、意味の読めない文字の塊になる地点で止まります。PaddleOCRはそこを支える層です。100を超える言語の文字を認識し、文書向けの処理はレイアウトが持っていた意味を保ちます。見出しは見出しのまま、表はセルの位置を伴う表として出力され、数式や印影も扱えます。出力形式はMarkdownまたはJSONです。モデルはCPUや端末側でも動く大きさに抑えられており、大量の文書をまとめて処理する費用が現実的になります。難点は土台となるフレームワークです。これはPaddlePaddleのプロジェクトであり、導入するとはそれを導入することを意味します。

PaddleOCRで何ができますか?

  • 文字だけでなく構造を残す — 文書向けの処理は、見出し、表、読み順、数式が保たれた形でMarkdownやJSONを出力します。検索で使える出力になるかどうかはここで決まります。
  • 必要ならセル単位の座標まで得る — 構造解析の処理は、各表セルや文字ブロックがページのどこにあったかを返します。抽出した数値を元の場所まで遡って確認できます。
  • モデルを切り替えずに100言語を読む — 1つの認識モデルが中国語、英語、日本語とラテン文字系の多数の言語を覆うため、言語の混在した文書群でも振り分けの処理が要りません。
  • 手持ちの機材で動かす — モデルは数メガバイト級のものから揃い、GPUや各種アクセラレータだけでなく通常のCPUでも推論できます。
  • 扱いにくい原本に対応する — 傾き、歪み、照明の悪さ、画面を撮影したページが明示的な対象で、印影や産業用の刻印なども想定されています。
  • すでに使っている道具に組み込む — DifyやRAGFlowなどが文書処理の層としてこれを使っており、出力形式はそれらの処理がそのまま受け取れるものです。

PaddleOCRを選ぶ前に

  • 土台はPaddlePaddleという深層学習フレームワークで、既に導入している環境は多くありません。導入作業は小さなパッケージ1つではなく、版と対応機材の確認を伴うフレームワークの導入になります。
  • モデルの世代交代が速く、名前も似ています。認識モデルの複数世代と2種類の文書処理が同時に現役のため、新しさではなく用途で選ぶ必要があります。

よくある質問

PaddleOCRは商用利用できますか?

PaddleOCRはApache-2.0ライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

PaddleOCRはどの形で使えますか?

PaddleOCRはローカル実行・セルフホストの形で利用できます。

ドキュメント

PaddlePaddle/PaddleOCR のREADMEより転載(Apache-2.0)。 原文を読む ↗

English | 简体中文 | 繁體中文 | 日本語 | 한국어 | Français | Русский | Español | العربية

PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown) with industry-leading accuracy. With 70k+ Stars and trusted by top-tier projects like Dify, RAGFlow, and Cherry Studio, PaddleOCR is the bedrock for building intelligent RAG and Agentic applications.

🚀 Key Features

📄 Intelligent Document Parsing (LLM-Ready)

Transforming messy visuals into structured data for the LLM era.

  • SOTA Document VLM: Featuring PaddleOCR-VL-1.6 (0.9B), the industry’s leading lightweight vision-language model for document parsing. It achieves 96.3% accuracy on OmniDocBench v1.6, leads in text, formula, and table recognition, and shows significantly enhanced capabilities in ancient documents, rare characters, seals, and charts, with structured outputs in Markdown and JSON formats.
  • Structure-Aware Conversion: Powered by PP-StructureV3, seamlessly convert complex PDFs and images into Markdown or JSON. Unlike the PaddleOCR-VL series models, it provides more fine-grained coordinate information, including table cell coordinates, text coordinates, and more.
  • Production-Ready Efficiency: Achieve commercial-grade accuracy with an ultra-small footprint. Outperforms numerous closed-source solutions in public benchmarks while remaining resource-efficient for edge/cloud deployment.

🔍 Universal Text Recognition (Scene OCR)

The global gold standard for high-speed, multilingual text spotting.

  • 100+ Languages Supported: Native recognition for a vast global library. PP-OCRv6 supports 50 languages with a single unified model (Chinese, English, Japanese, and 46 Latin-script languages) — no model switching needed for multilingual documents.
  • Complex Element Mastery: Beyond standard text recognition, we support natural scene text spotting across a wide range of environments, including IDs, street views, books, and industrial components
  • Performance Leap: PP-OCRv6 achieves +4.6% detection and +5.1% recognition accuracy over PP-OCRv5, surpassing mainstream Vision-Language Models. 5.2× CPU inference speedup end-to-end.

🛠️ Developer-Centric Ecosystem

  • Seamless Integration: The premier choice for the AI Agent ecosystem—deeply integrated with Dify, RAGFlow, Pathway, and Cherry Studio.
  • LLM Data Flywheel: A complete pipeline to build high-quality datasets, providing a sustainable “Data Engine” for fine-tuning Large Language Models.
  • One-Click Deployment: Supports various hardware backends (NVIDIA GPU, Intel CPU, Kunlunxin XPU, and diverse AI Accelerators).

📣 Recent updates

🔥 2026.07.22: HPD-Parsing is now available

  • HPD-Parsing is a lightweight vision-language model designed for high-throughput document parsing. It adopts a hierarchical parallel decoding paradigm and Progressive Multi-Token Prediction (P-MTP), achieving a peak throughput of 4,752 tokens/s on public benchmarks while maintaining competitive parsing accuracy.
  • HPD-Parsing supports both OpenAI-compatible serving and local inference through a customized vLLM runtime, making it suitable for document parsing scenarios with high demands on inference efficiency and deployment throughput.
  • See the HPD-Parsing usage tutorial for environment setup, serving, and local inference instructions.
  • PP-OCRv6 highlights:

    • Accuracy boost: Medium tier achieves +4.6% detection and +5.1% recognition over PP-OCRv5_server, surpassing mainstream VLMs (Qwen3-VL-235B, GPT-5.5) with only 34.5M parameters.
    • 50 languages unified: Single model covers Chinese, English, Japanese, and 46 Latin-script languages — no model switching needed.
    • Specialized scenarios: Major improvements in digital displays, dot-matrix characters, tire prints, and industrial text recognition.
    • Faster inference: 5.2× CPU speedup (OpenVINO), 6.1× on Apple M4 (tiny), 0.13s on A100 GPU.
    • Three tiers for all scenarios: tiny (1.5M) / small (7.7M) / medium (34.5M) for edge, mobile, and server deployment.
    • Model availability: All models are available on HuggingFace and ModelScope.
  • PaddleOCR-VL-1.6 highlights:

    • New SOTA Accuracy: Achieves over 96.3% on OmniDocBench v1.6, also sets new SOTA on OmniDocBench v1.5 and Real5-OmniDocBench, leading both open-source and proprietary solutions in text, formula, and table recognition.
    • Comprehensive Capability Upgrade: Significant improvements in table, ancient document, and rare character recognition, with notably enhanced seal recognition, spotting, and chart understanding across multiple scenarios.
    • Seamless Migration: Model architecture is fully consistent with PaddleOCR-VL-1.5, enabling zero-cost adaptation—swap and go.
    • Try it now: Available on HuggingFace or our Official Website.
  • Flexible inference backends: Seamlessly switch between Paddle static graph, Paddle dynamic graph, or Transformers. PaddleOCR is now deeply integrated with the Hugging Face ecosystem, and 20 major models support Transformers as the inference backend.
  • Office documents to Markdown: Convert common document formats such as Word, Excel, and PowerPoint into Markdown.
  • DOCX export for parsed results: The PaddleOCR-VL series, PP-StructureV3, and PP-DocTranslation now support exporting parsed results to DOCX for convenient viewing and editing in Microsoft Word.
  • Official browser inference SDK: Released PaddleOCR.js, the official browser inference SDK that supports running PP-OCRv5 directly in the browser.
  • PaddleOCR-VL-1.5 (SOTA 0.9B VLM): Our latest flagship model for document parsing is now live!
    • 94.5% Accuracy on OmniDocBench: Surpassing top-tier general large models and specialized document parsers.
    • Real-World Robustness: First to introduce the PP-DocLayoutV3 algorithm for irregular shape positioning, mastering 5 tough scenarios: Skew, Warping, Scanning, Illumination, and Screen Photography.
    • Capability Expansion: Now supports Seal Recognition, Text Spotting, and expands to 111 languages (including China’s Tibetan script and Bengali).
    • Long Document Mastery: Supports automatic cross-page table merging and hierarchical heading identification.
    • Try it now: Available on HuggingFace or our Official Website.
  • Released PaddleOCR-VL:

    • Model Introduction:

      • PaddleOCR-VL is a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-language model (VLM) that integrates a NaViT-style dynamic resolution visual encoder with the ERNIE-4.5-0.3B language model to enable accurate element recognition. This innovative model efficiently supports 109 languages and excels in recognizing complex elements (e.g., text, tables, formulas, and charts), while maintaining minimal resource consumption. Through comprehensive evaluations on widely used public benchmarks and in-house benchmarks, PaddleOCR-VL achieves SOTA performance in both page-level document parsing and element-level recognition. It significantly outperforms existing solutions, exhibits strong competitiveness against top-tier VLMs, and delivers fast inference speeds. These strengths make it highly suitable for practical deployment in real-world scenarios. The model has been released on HuggingFace. Everyone is welcome to download and use it! More introduction information can be found in PaddleOCR-VL.
    • Core Features:

      • Compact yet Powerful VLM Architecture: We present a novel vision-language model that is specifically designed for resource-efficient inference, achieving outstanding performance in element recognition. By integrating a NaViT-style dynamic high-resolution visual encoder with the lightweight ERNIE-4.5-0.3B language model, we significantly enhance the model’s recognition capabilities and decoding efficiency. This integration maintains high accuracy while reducing computational demands, making it well-suited for efficient and practical document processing applications.
      • SOTA Performance on Document Parsing: PaddleOCR-VL achieves state-of-the-art performance in both page-level document parsing and element-level recognition. It significantly outperforms existing pipeline-based solutions and exhibiting strong competitiveness against leading vision-language models (VLMs) in document parsing. Moreover, it excels in recognizing complex document elements, such as text, tables, formulas, and charts, making it suitable for a wide range of challenging content types, including handwritten text and historical documents. This makes it highly versatile and suitable for a wide range of document types and scenarios.
      • Multilingual Support: PaddleOCR-VL Supports 109 languages, covering major global languages, including but not limited to Chinese, English, Japanese, Latin, and Korean, as well as languages with different scripts and structures, such as Russian (Cyrillic script), Arabic, Hindi (Devanagari script), and Thai. This broad language coverage substantially enhances the applicability of our system to multilingual and globalized document processing scenarios.
  • Released PP-OCRv5 Multilingual Recognition Model:

    • Improved the accuracy and coverage of Latin script recognition; added support for Cyrillic, Arabic, Devanagari, Telugu, Tamil, and other language systems, covering recognition of 109 languages. The model has only 2M parameters, and the accuracy of some models has increased by over 40% compared to the previous generation.
  • Significant Model Additions:

    • Introduced training, inference, and deployment for PP-OCRv5 recognition models in English, Thai, and Greek. The PP-OCRv5 English model delivers an 11% improvement in English scenarios compared to the main PP-OCRv5 model, with the Thai and Greek recognition models achieving accuracies of 82.68% and 89.28%, respectively.
  • Deployment Capability Upgrades:

    • Full support for PaddlePaddle framework versions 3.1.0 and 3.1.1.
    • Comprehensive upgrade of the PP-OCRv5 C++ local deployment solution, now supporting both Linux and Windows, with feature parity and identical accuracy to the Python implementation.
    • High-performance inference now supports CUDA 12, and inference can be performed using either the Paddle Inference or ONNX Runtime backends.
    • The high-stability service-oriented deployment solution is now fully open-sourced, allowing users to customize Docker images and SDKs as required.
    • The high-stability service-oriented deployment solution also supports invocation via manually constructed HTTP requests, enabling client-side code development in any programming language.
  • Benchmark Support:

    • All production lines now support fine-grained benchmarking, enabling measurement of end-to-end inference time as well as per-layer and per-module latency data to assist with performance analysis. Here’s how to set up and use the benchmark feature.
    • Documentation has been updated to include key metrics for commonly used configurations on mainstream hardware, such as inference latency and memory usage, providing deployment references for users.
  • Bug Fixes:

    • Resolved the issue of failed log saving during model training.
    • Upgraded the data augmentation component for formula models for compatibility with newer versions of the albumentations dependency, and fixed deadlock warnings when using the tokenizers package in multi-process scenarios.
    • Fixed inconsistencies in switch behaviors (e.g., use_chart_parsing) in the PP-StructureV3 configuration files compared to other pipelines.
  • Other Enhancements:

    • Separated core and optional dependencies. Only minimal core dependencies are required for basic text recognition; additional dependencies for document parsing and information extraction can be installed as needed.
    • Enabled support for NVIDIA RTX 50 series graphics cards on Windows; users can refer to the installation guide for the corresponding PaddlePaddle framework versions.
    • PP-OCR series models now support returning single-character coordinates.
    • Added AIStudio, ModelScope, and other model download sources, allowing users to specify the source for model downloads.
    • Added support for chart-to-table conversion via the PP-Chart2Table module.
    • Optimized documentation descriptions to improve usability.

History Log

🚀 Quick Start

Step 1: Try Online

PaddleOCR official website provides interactive Experience Center and APIs—no setup required, just one click to experience.

👉 Visit Official Website

Step 2: Local Deployment

For local usage, please refer to the following documentation based on your needs:

🧩 More Features

🔄 Quick Overview of Execution Results

PP-OCRv5

PP-StructureV3

PaddleOCR-VL

✨ Stay Tuned

⭐ Star this repository to keep up with exciting updates and new releases, including powerful OCR and document parsing capabilities! ⭐

👩‍👩‍👧‍👦 Community

PaddlePaddle WeChat official accountJoin the tech discussion group

このREADMEは一部を省略しています。全文はGitHubにあります。 原文を読む ↗

PaddleOCR
AIに聞く
GitHub