
Docling
PDF・Officeファイル・スキャンを、外部に送らずに1つの構造化表現へ変換するツール
Doclingとは
テキスト抽出との違いは構造が残る点です。ページの版面、読み順、表のセル、コードブロック、数式が、平坦な文字列ではなく単一の文書表現の一部として出てきます。報告書の中の表がエージェントに届く時点でも表のままかどうかは、ここで決まります。処理は完全にローカルで実行でき、外部と切り離された環境でも動きます。その忠実さの代償は計算量です。各ページに機械学習モデルを走らせるため、その場で呼ぶ高速な変換器ではなく、見積もっておくべき処理として扱う必要があります。
Doclingで何ができますか?
- 入力が何でも表現は1つ — 入力が何であれ、出てくるのは単一の文書形式です。後段のコードが、元がPDFか表計算かメールかで分岐する必要がありません。
- テキスト化で失われる構造を残す — ページの版面、読み順、表の構造、コードブロック、数式、画像の分類が復元されます。グラフは説明を添えた表またはコードに変換されます。
- 扱いにくい形式も読む — PDFとOfficeのほか、HTML、EPUB、画像、LaTeX、メール、音声認識を通した音声に対応し、XBRLの財務報告やUSPTOの特許といった専用スキーマも扱えます。
- スキャン文書を外部サービスなしで処理する — スキャンされたPDFや画像向けのOCRを内蔵し、視覚言語モデルにも対応します。文字が画像になっているページが単に飛ばされることがありません。
- 結果をそのままRAG構成へ渡す — LangChain、LlamaIndex、CrewAI、Haystackとの差し込み型の統合があります。変換をライブラリ呼び出しではなくサービスとして扱いたい場合のAPIサーバーモードも用意されています。
Doclingを選ぶ前に
- 版面と表構造の機械学習モデルを各ページに走らせるため、変換の処理量は設備の問題になります。リクエストの経路上で呼ぶ高速な抽出器ではなく、計画して回す処理です。
スター推移
8月21日〜8月28日 · +335
よくある質問
Doclingは商用利用できますか?
DoclingはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
Doclingはどの形で使えますか?
Doclingはローカル実行・セルフホストの形で利用できます。
ドキュメント
docling-project/docling のREADMEより転載(MIT)。 原文を読む ↗
Docling
What is Docling ?
Docling simplifies document processing by parsing diverse formats — including advanced PDF understanding — and providing seamless integrations with the generative AI ecosystem.
Features
- 🗂️ Parsing of multiple document formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, WAV, MP3, WebVTT, Box Notes, email formats (EML, MSG), images (PNG, TIFF, JPEG, …), LaTeX, DocLang, plain text, and more
- 📑 Advanced PDF understanding incl. page layout, reading order, table structure, code, formulas, image classification, and more
- 🧬 A unified, expressive DoclingDocument representation format
- ↪️ Various export formats and options, including Markdown, HTML, WebVTT, DocLang, DocTags and lossless JSON
- 📜 Support for several application-specific XML schemas including DocLang, USPTO patents, JATS articles, and XBRL financial reports.
- 🔒 Local execution capabilities for sensitive data and air-gapped environments
- 🤖 Plug-and-play integrations incl. LangChain, LlamaIndex, Crew AI & Haystack for agentic AI
- 🔍 Extensive OCR support for scanned PDFs and images
- 👓 Support for several Visual Language Models, such as (GraniteDocling)
- 🎙️ Audio support with Automatic Speech Recognition (ASR) models
- 🔌 Connect to any agent using the MCP server
- 🌐 Run Docling as a service with the API server (docling-serve)
- 💻 Simple and convenient CLI
What’s new
- 🎬 Parsing of video files (MP4, AVI, MOV, MKV, and WebM) with an ASR transcript and representative keyframes
- 📄 Parsing of ODF (OpenDocument Format) files for text documents (
.odt), spreadsheets (.ods), and presentations (.odp) - 💼 Parsing of XBRL (eXtensible Business Reporting Language) documents for financial reports
- 📧 Parsing of email files (
.eml,.msg) - 📚 Parsing of EPUB (Electronic Publication) files for e-books
- 📝 Parsing of plain-text files (
.txt,.text) and Markdown supersets (.qmd,.Rmd) - 📊 Chart understanding (Barchart, Piechart, LinePlot): convert them into tables or code and add detailed descriptions
Coming soon
- 📝 Metadata extraction, including title, authors, references & language
- 📝 Complex chemistry understanding (Molecular structures)
Quickstart
1. Install
pip install docling
Note: Python 3.9 support was dropped in docling version 2.70.0. Please use Python 3.10 or higher.
Works on macOS, Linux and Windows environments for both x86_64 and arm64 architectures.
More detailed installation instructions are available in the docs.
2. Convert a document (CLI)
docling https://arxiv.org/pdf/2206.01062
This generates a .md file in the current directory containing structured document content.
You can also use 🥚GraniteDocling and other VLMs via Docling CLI:
docling --pipeline vlm --vlm-model granite_docling https://arxiv.org/pdf/2206.01062
3. Python usage (recommended)
from docling.document_converter import DocumentConverter
source = "https://arxiv.org/pdf/2408.09869" # a document via a local path or URL
converter = DocumentConverter()
result = converter.convert(source)
print(result.document.export_to_markdown()) # output: "## Docling Technical Report[...]"
More advanced usage and configuration options.
Documentation
Check out Docling’s documentation for details on installation, usage, concepts, recipes, extensions, and more.
Examples
Go hands-on with our examples, demonstrating how to address different application use cases with Docling.
Integrations
To further accelerate your AI application development, check out Docling’s native integrations with popular frameworks and tools.
Get help and support
Please feel free to connect with us using the discussion section.
Technical report
For more details on Docling’s inner workings, check out the Docling Technical Report.
References
If you use Docling in your projects, please consider citing the following:
@techreport{Docling,
author = {Deep Search Team},
month = {8},
title = {Docling Technical Report},
url = {https://arxiv.org/abs/2408.09869},
eprint = {2408.09869},
doi = {10.48550/arXiv.2408.09869},
version = {1.0.0},
year = {2024}
}
LF AI & Data
Docling is hosted as a project in the LF AI & Data Foundation.
IBM ❤️ Open Source AI
The project was started by the AI for knowledge team at IBM Research Zurich.