← Back to all projects

Docling

Converts PDFs, Office files and scans into one structured representation, without sending them anywhere

OfficialMIT
Stars
65.7k
Forks
4.7k
Open issues
979
Last commit
28 Aug 2026

What is Docling?

The difference from a text extractor is that structure survives: page layout, reading order, table cells, code blocks and formulas come through as parts of a single document representation rather than as a flattened string, which is what decides whether a table in a report is still a table by the time an agent reads it. Everything can run locally, including in air-gapped environments. The cost of that fidelity is compute — machine-learning models run over each page, so this is a pipeline to budget for, not a fast converter to call inline.

What can you do with Docling?

  • One representation for every input — Whatever went in, what comes out is a single document format, so downstream code does not branch on whether the source was a PDF, a spreadsheet or an email.
  • Keep the structure a text dump destroys — Page layout, reading order, table structure, code blocks, formulas and image classification are all recovered, and charts are converted into tables or code with a description attached.
  • Read the awkward formats too — Beyond PDF and Office there is HTML, EPUB, images, LaTeX, email, audio via speech recognition, and specialised schemas such as XBRL financial filings and USPTO patents.
  • Handle scans without an external service — OCR for scanned PDFs and images is included, along with support for vision-language models, so pages that are pictures of text are not simply skipped.
  • Hand the result straight to a RAG stack — Plug-in integrations exist for LangChain, LlamaIndex, CrewAI and Haystack, and an API server mode is available when conversion should be a service rather than a library call.

Before you choose Docling

  • Machine-learning models for layout and table structure run over each page, so conversion throughput is a capacity question — this is a pipeline to schedule, not a fast extractor to call in a request path.

Star history

21 Aug to 28 Aug · +335

65.4k65.7k

Frequently asked questions

Is Docling free for commercial use?

Docling is released under the MIT licence — OSI-approved open source, which permits commercial use.

How can Docling be deployed?

Docling is available as Runs locally / Self-hosted.

Documentation

Reproduced from the docling-project/docling README, published under MIT. Read the original ↗

Docling

Docling Actor Chat with Dosu OpenSSF Best Practices

What is Docling ?

Docling simplifies document processing by parsing diverse formats — including advanced PDF understanding — and providing seamless integrations with the generative AI ecosystem.

Features

  • 🗂️ Parsing of multiple document formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, WAV, MP3, WebVTT, Box Notes, email formats (EML, MSG), images (PNG, TIFF, JPEG, …), LaTeX, DocLang, plain text, and more
  • 📑 Advanced PDF understanding incl. page layout, reading order, table structure, code, formulas, image classification, and more
  • 🧬 A unified, expressive DoclingDocument representation format
  • ↪️ Various export formats and options, including Markdown, HTML, WebVTT, DocLang, DocTags and lossless JSON
  • 📜 Support for several application-specific XML schemas including DocLang, USPTO patents, JATS articles, and XBRL financial reports.
  • 🔒 Local execution capabilities for sensitive data and air-gapped environments
  • 🤖 Plug-and-play integrations incl. LangChain, LlamaIndex, Crew AI & Haystack for agentic AI
  • 🔍 Extensive OCR support for scanned PDFs and images
  • 👓 Support for several Visual Language Models, such as (GraniteDocling)
  • 🎙️ Audio support with Automatic Speech Recognition (ASR) models
  • 🔌 Connect to any agent using the MCP server
  • 🌐 Run Docling as a service with the API server (docling-serve)
  • 💻 Simple and convenient CLI

What’s new

  • 🎬 Parsing of video files (MP4, AVI, MOV, MKV, and WebM) with an ASR transcript and representative keyframes
  • 📄 Parsing of ODF (OpenDocument Format) files for text documents (.odt), spreadsheets (.ods), and presentations (.odp)
  • 💼 Parsing of XBRL (eXtensible Business Reporting Language) documents for financial reports
  • 📧 Parsing of email files (.eml, .msg)
  • 📚 Parsing of EPUB (Electronic Publication) files for e-books
  • 📝 Parsing of plain-text files (.txt, .text) and Markdown supersets (.qmd, .Rmd)
  • 📊 Chart understanding (Barchart, Piechart, LinePlot): convert them into tables or code and add detailed descriptions

Coming soon

  • 📝 Metadata extraction, including title, authors, references & language
  • 📝 Complex chemistry understanding (Molecular structures)

Quickstart

1. Install

pip install docling

Note: Python 3.9 support was dropped in docling version 2.70.0. Please use Python 3.10 or higher.

Works on macOS, Linux and Windows environments for both x86_64 and arm64 architectures.

More detailed installation instructions are available in the docs.

2. Convert a document (CLI)

docling https://arxiv.org/pdf/2206.01062

This generates a .md file in the current directory containing structured document content.

You can also use 🥚GraniteDocling and other VLMs via Docling CLI:

docling --pipeline vlm --vlm-model granite_docling https://arxiv.org/pdf/2206.01062
from docling.document_converter import DocumentConverter

source = "https://arxiv.org/pdf/2408.09869"  # a document via a local path or URL
converter = DocumentConverter()
result = converter.convert(source)
print(result.document.export_to_markdown())  # output: "## Docling Technical Report[...]"

More advanced usage and configuration options.

Documentation

Check out Docling’s documentation for details on installation, usage, concepts, recipes, extensions, and more.

Examples

Go hands-on with our examples, demonstrating how to address different application use cases with Docling.

Integrations

To further accelerate your AI application development, check out Docling’s native integrations with popular frameworks and tools.

Get help and support

Please feel free to connect with us using the discussion section.

Technical report

For more details on Docling’s inner workings, check out the Docling Technical Report.

References

If you use Docling in your projects, please consider citing the following:

@techreport{Docling,
  author = {Deep Search Team},
  month = {8},
  title = {Docling Technical Report},
  url = {https://arxiv.org/abs/2408.09869},
  eprint = {2408.09869},
  doi = {10.48550/arXiv.2408.09869},
  version = {1.0.0},
  year = {2024}
}

LF AI & Data

Docling is hosted as a project in the LF AI & Data Foundation.

IBM ❤️ Open Source AI

The project was started by the AI for knowledge team at IBM Research Zurich.