PaddleOCR

Package

Open-source OCR toolkit you self-host, with formula-to-LaTeX and PDF parsing

Price
Free, open source
Access
None, runs locally

About

Apache-2.0 OCR toolkit from the PaddlePaddle team that runs on your own machines. PP-OCRv6 reads 50 languages with one model (100+ in all), a formula pipeline turns equation images into LaTeX, and PP-StructureV3 and PaddleOCR-VL convert PDFs to Markdown or JSON. Needs PaddlePaddle.

What you can do with it

  • Convert images of math formulas into LaTeX on your own server
  • Read text from photos and scans in 100+ languages without a cloud API
  • Turn PDFs into Markdown or JSON with tables and formulas for RAG

Get started

  1. Install PaddlePaddle for your hardware, then pip install paddleocr
  2. Run paddleocr formula_recognition_pipeline -i equation.png

Example

from paddleocr import FormulaRecognitionPipeline

pipeline = FormulaRecognitionPipeline()
for res in pipeline.predict("./general_formula_recognition_001.png"):
    res.print()  # LaTeX for each recognized formula
    res.save_to_json(save_path="output")

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python, JavaScript/TypeScript
MCP server
Local

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .