PaddleOCR
PackageOpen-source OCR toolkit you self-host, with formula-to-LaTeX and PDF parsing
- Price
- Free, open source
- Access
- None, runs locally
About
Apache-2.0 OCR toolkit from the PaddlePaddle team that runs on your own machines. PP-OCRv6 reads 50 languages with one model (100+ in all), a formula pipeline turns equation images into LaTeX, and PP-StructureV3 and PaddleOCR-VL convert PDFs to Markdown or JSON. Needs PaddlePaddle.
What you can do with it
- Convert images of math formulas into LaTeX on your own server
- Read text from photos and scans in 100+ languages without a cloud API
- Turn PDFs into Markdown or JSON with tables and formulas for RAG
Get started
- Install PaddlePaddle for your hardware, then pip install paddleocr
- Run paddleocr formula_recognition_pipeline -i equation.png
Example
from paddleocr import FormulaRecognitionPipeline
pipeline = FormulaRecognitionPipeline()
for res in pipeline.predict("./general_formula_recognition_001.png"):
res.print() # LaTeX for each recognized formula
res.save_to_json(save_path="output")Details
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- Official SDKs
- Python, JavaScript/TypeScript
- MCP server
- Local