Docling

Package

Open-source Python library that parses PDFs, Office files and images

Price
Free, open source
Access
None, runs locally

About

MIT-licensed Python library and CLI that turns PDFs, Office files, HTML, images and audio into Markdown or lossless JSON, keeping layout, reading order and tables, with OCR for scans. It runs locally, even air-gapped, and needs Python 3.10 or later.

What you can do with it

  • Convert PDFs, DOCX, PPTX and XLSX files into clean Markdown or JSON for RAG
  • Extract tables, reading order and page layout from complex PDF reports
  • OCR scanned PDFs and images on your own machine without sending data out

Get started

  1. Install with pip install docling (Python 3.10+)
  2. Pass a local path or URL to DocumentConverter

Example

pip install docling
python - <<'EOF'
from docling.document_converter import DocumentConverter

result = DocumentConverter().convert("https://arxiv.org/pdf/2408.09869")
print(result.document.export_to_markdown())
EOF

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python
MCP server
Local

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .