Amazon Textract
APIAWS API that reads text, forms, tables and invoices from scanned documents
- Access
- AWS credentials
About
AWS service that detects printed and handwritten text and analyzes forms, tables, invoices, receipts, US IDs and mortgage packages. Synchronous calls take one page; multipage PDFs go through S3 with async jobs. Text detection covers English and five other European languages.
What you can do with it
- Detect printed and handwritten text in scanned PDFs and images
- Extract forms, tables and answers to natural-language queries from documents
- Read vendor, totals and line items from invoices and receipts
Get started
- Create an IAM user with Textract access and an access key
- Configure the AWS CLI with that profile
- Upload the document to an S3 bucket in the same region
Example
aws textract detect-document-text \
--document '{"S3Object":{"Bucket":"my-bucket","Name":"scan.png"}}' \
--region us-east-1
# Invoices and receipts
aws textract analyze-expense \
--document '{"S3Object":{"Bucket":"my-bucket","Name":"invoice.pdf"}}' \
--region us-east-1Details
- Hosting
- Hosted service
- Available in
- Worldwide
- Official SDKs
- Python, JavaScript/TypeScript, Java, Go, Ruby, PHP, C#, Kotlin, Rust, Swift
- MCP server
- None
- Works with
- AWS, S3