Tesseract OCR

Package

Open-source OCR engine you self-host that reads 100+ languages from images

Price
Free, open source
Access
None, runs locally

About

Tesseract is the Apache-2.0 OCR engine and CLI behind many OCR tools. It reads text in 100+ languages from images on your own machine and writes plain text, hOCR, PDF, TSV, ALTO or PAGE. Python and JavaScript use community wrappers; image quality matters, so preprocess scans.

What you can do with it

  • Extract text from scanned images on your own server, with no cloud API
  • Turn scans into searchable PDF or hOCR output with word positions
  • OCR documents in 100+ languages and scripts from the command line

Get started

  1. Install it, e.g. brew install tesseract or apt install tesseract-ocr
  2. Add language data such as tesseract-ocr-deu for other languages

Example

# Plain text in output.txt
tesseract scan.png output -l eng

# Searchable PDF in output.pdf, English and German
tesseract scan.png output -l eng+deu pdf

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .