OCRmyPDF

Package

Self-hosted CLI that adds a searchable OCR text layer to scanned PDFs

Price
Free, open source
Access
None, runs locally

About

OCRmyPDF runs Tesseract on scanned PDFs on your own machine and writes a searchable PDF/A with the text placed under each page image, so it can be searched and copied. It deskews and rotates pages, optimizes images and uses all CPU cores. CLI, Python API and Docker images.

What you can do with it

  • Make scanned PDFs searchable and copyable without sending them to a cloud API
  • Batch-convert an archive of scans into PDF/A for long-term storage
  • Deskew and rotate crooked scanned pages while adding the OCR text

Get started

  1. Install it, e.g. apt install ocrmypdf or brew install ocrmypdf
  2. Run ocrmypdf input.pdf output.pdf

Example

ocrmypdf -l eng --rotate-pages --deskew input_scanned.pdf output_searchable.pdf

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .