Tesseract OCR
PackageOpen-source OCR engine you self-host that reads 100+ languages from images
- Price
- Free, open source
- Access
- None, runs locally
About
Tesseract is the Apache-2.0 OCR engine and CLI behind many OCR tools. It reads text in 100+ languages from images on your own machine and writes plain text, hOCR, PDF, TSV, ALTO or PAGE. Python and JavaScript use community wrappers; image quality matters, so preprocess scans.
What you can do with it
- Extract text from scanned images on your own server, with no cloud API
- Turn scans into searchable PDF or hOCR output with word positions
- OCR documents in 100+ languages and scripts from the command line
Get started
- Install it, e.g. brew install tesseract or apt install tesseract-ocr
- Add language data such as tesseract-ocr-deu for other languages
Example
# Plain text in output.txt
tesseract scan.png output -l eng
# Searchable PDF in output.pdf, English and German
tesseract scan.png output -l eng+deu pdfDetails
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- MCP server
- None