PaddleOCR
PaketOpen-source OCR toolkit you self-host, with formula-to-LaTeX and PDF parsing
- Fiyat
- Free, open source
- Erişim
- None, runs locally
Hakkında
Apache-2.0 OCR toolkit from the PaddlePaddle team that runs on your own machines. PP-OCRv6 reads 50 languages with one model (100+ in all), a formula pipeline turns equation images into LaTeX, and PP-StructureV3 and PaddleOCR-VL convert PDFs to Markdown or JSON. Needs PaddlePaddle.
Neler yapabilirsin
- Convert images of math formulas into LaTeX on your own server
- Read text from photos and scans in 100+ languages without a cloud API
- Turn PDFs into Markdown or JSON with tables and formulas for RAG
Başlarken
- Install PaddlePaddle for your hardware, then pip install paddleocr
- Run paddleocr formula_recognition_pipeline -i equation.png
Örnek kod
from paddleocr import FormulaRecognitionPipeline
pipeline = FormulaRecognitionPipeline()
for res in pipeline.predict("./general_formula_recognition_001.png"):
res.print() # LaTeX for each recognized formula
res.save_to_json(save_path="output")Ayrıntılar
- Barındırma
- Kendi sunucunda, Kendi bilgisayarında
- Kullanılabildiği yerler
- Tüm dünya
- Resmi SDK'lar
- Python, JavaScript/TypeScript
- MCP sunucusu
- Yerel sunucu