OCRmyPDF
PackageSelf-hosted CLI that adds a searchable OCR text layer to scanned PDFs
- Price
- Free, open source
- Access
- None, runs locally
About
OCRmyPDF runs Tesseract on scanned PDFs on your own machine and writes a searchable PDF/A with the text placed under each page image, so it can be searched and copied. It deskews and rotates pages, optimizes images and uses all CPU cores. CLI, Python API and Docker images.
What you can do with it
- Make scanned PDFs searchable and copyable without sending them to a cloud API
- Batch-convert an archive of scans into PDF/A for long-term storage
- Deskew and rotate crooked scanned pages while adding the OCR text
Get started
- Install it, e.g. apt install ocrmypdf or brew install ocrmypdf
- Run ocrmypdf input.pdf output.pdf
Example
ocrmypdf -l eng --rotate-pages --deskew input_scanned.pdf output_searchable.pdfDetails
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- Official SDKs
- Python
- MCP server
- None