Mistral OCR

Model

OCR model API that turns PDFs, Office files and images into markdown

Price
Free tier, then paid
Access
API key

About

Mistral's OCR model, served through the Mistral API, reads PDFs, Office files and images and returns markdown per page with tables, bounding boxes and confidence scores. Annotations extract JSON in your schema. Files are limited to 50 MB and 1,000 pages.

What you can do with it

  • Convert scanned or digital PDFs into markdown with tables and images kept in place
  • Extract fields from documents into JSON with a schema you define
  • Batch-process large document sets at half price through the Batch API

Get started

  1. Create an API key in Mistral Studio
  2. Pass a document URL or upload the file first

Example

curl https://api.mistral.ai/v1/ocr \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_API_KEY" \
  -d '{
    "model": "mistral-ocr-4-1",
    "document": {
      "type": "document_url",
      "document_url": "https://arxiv.org/pdf/2201.04234"
    },
    "table_format": "html"
  }'

Details

Hosting
Hosted service, Self-hosted
Available in
Worldwide
Official SDKs
Python, JavaScript/TypeScript
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .