Docling
PackageOpen-source Python library that parses PDFs, Office files and images
- Price
- Free, open source
- Access
- None, runs locally
About
MIT-licensed Python library and CLI that turns PDFs, Office files, HTML, images and audio into Markdown or lossless JSON, keeping layout, reading order and tables, with OCR for scans. It runs locally, even air-gapped, and needs Python 3.10 or later.
What you can do with it
- Convert PDFs, DOCX, PPTX and XLSX files into clean Markdown or JSON for RAG
- Extract tables, reading order and page layout from complex PDF reports
- OCR scanned PDFs and images on your own machine without sending data out
Get started
- Install with pip install docling (Python 3.10+)
- Pass a local path or URL to DocumentConverter
Example
pip install docling
python - <<'EOF'
from docling.document_converter import DocumentConverter
result = DocumentConverter().convert("https://arxiv.org/pdf/2408.09869")
print(result.document.export_to_markdown())
EOFDetails
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- Official SDKs
- Python
- MCP server
- Local