Claude Opus 5.5 is the strongest model for reading scans and images. MiMo-V2.6-Pro extracts text for 7% of the cost, and DeepSeek V4.1 Flash answers fastest.
Claude Opus 5.5 is the strongest model for reading scans and images. MiMo-V2.6-Pro extracts text for 7% of the cost, and DeepSeek V4.1 Flash answers fastest.
27 models · Updated September 2026 · Data from Artificial Analysis
Claude Opus 5.5
Anthropic · High effort
The highest score of any model that accepts images at this setting.
MiMo-V2.6-Pro
Xiaomi · Standard effort
86% of Claude Opus 5.5's score for 7% of the cost.
DeepSeek V4.1 Flash
DeepSeek · Max effort
The quickest full answer from a model scoring 38 or more.
Over 45 s at every tested setting
Claude Opus 5.5
Anthropic · High effort
The highest score of any model that accepts images at this setting.
MiMo-V2.6-Pro
Xiaomi · Standard effort
86% of Claude Opus 5.5's score for 7% of the cost.
DeepSeek V4.1 Flash
DeepSeek · Max effort
The quickest full answer from a model scoring 38 or more.
Each job favors a different model. Pick the one closest to yours.
Messy scans, handwriting and tables
Use
Claude Opus 5.5
Anthropic · High effort
The strongest reasoning of any model that reads images, which helps most when the layout is complicated or the writing is hard to make out. $32.84 per 1,000 pages.
Messy scans, handwriting and tables
Use Claude Opus 5.5
The strongest reasoning of any model that reads images, which helps most when the layout is complicated or the writing is hard to make out. $32.84 per 1,000 pages.
Digitizing thousands of pages
Use MiMo-V2.6-Pro
$2.27 per 1,000 pages. About 61 seconds per page, fine for a batch job that runs overnight. MIT-licensed, so you can run it on your own servers if the documents are sensitive.
Receipts and IDs in a live app
Use DeepSeek V4.1 Flash
Reads a page in about 12 seconds for $3.67 per 1,000 pages.
Clean printed text on a tight budget
Use GPT-6 Luna
$0.78 per 1,000 pages, the cheapest paid option. Fine for clear print; check a sample before trusting it with handwriting.
Staying with Google
Use Gemini 3.8 Flash
Google's strongest current model for this, at 40.9 and $16.46 per 1,000 pages.
Whole PDFs, invoices and forms
Pulling specific fields out of multi-page documents is its own job, with its own ranking.
Understanding charts and photos
Explaining what an image shows needs more reasoning than reading its text.
Each dot is a model. Higher scores better, further left is cheaper, so the best deals sit top left. The cost axis is logarithmic: every step left is a big saving. Hover a dot for its numbers.
Sort by the pillar you care about, and set your monthly volume to see the bill. Costs include each model's thinking tokens at the setting shown.
Gemini 3.8 Flash. It scores 40.9 to Gemini 3.1 Pro's 29.7 and gives a full answer in 24 seconds instead of 35 seconds. Its list price is lower ($0.75/$3.75 per million tokens against $2/$12), but it thinks more, so 1,000 pages cost about $16.46 against $14.23.
Thinking helps with messy layouts, tables and handwriting. It helps much less with clean printed text, and every thinking token is billed.
For plain text extraction, start each model at its lowest reasoning setting, check a sample of pages, and raise the setting only where the output slips. Our costs assume the everyday setting, so a low setting will usually come in cheaper than shown.
Every model on this page accepts image input. We compare every model at the setting people run day to day: its strongest reasoning setting that gives a full answer within 45 seconds in Artificial Analysis's tests. Switch to Max effort above for the benchmark numbers.
Intelligence is the Artificial Analysis Intelligence Index at that setting. No public OCR benchmark covers every current model at everyday settings, so we rank on general intelligence and show cost per page and speed next to it. Accuracy on your own scans can differ, so test a sample before you commit.
Cost starts from a typical page: 3,000 input tokens for the image and your prompt, and 2,000 output tokens for the text. We scale the output by how much each model wrote, thinking included, in Artificial Analysis's tests, so a model that thinks a lot pays for it.
Speed is the median time to a full answer in the same tests.
Best value is the most intelligence per dollar and fastest is the quickest full answer, both among models scoring 38 or more (within 30% of the leader). Neither pick goes to a model its maker has replaced.
Claude Opus 5.5 is the strongest model that reads images at everyday settings (53.6 on the Artificial Analysis Intelligence Index), at about $32.84 per 1,000 pages. For volume, MiMo-V2.6-Pro costs $2.27 per 1,000 pages. No public OCR benchmark covers every model at these settings, so test a sample of your own documents.
Gemini 3.8 Flash. It scores 40.9 to Gemini 3.1 Pro's 29.7 and gives a full answer in 24 seconds instead of 35 seconds. Its list price is lower ($0.75/$3.75 per million tokens against $2/$12), but it thinks more, so 1,000 pages cost about $16.46 against $14.23.
GPT-6 Luna, at about $0.78 per 1,000 pages (list price $0.10/$0.50 per million tokens). It's fine for clean print. Qwen 2.5 VL 72B is free on OpenRouter, but it's older and rate-limited.
DeepSeek V4.1 Flash, at about 12 seconds per page at max effort. A lower reasoning setting can make it faster still on clean text.
Not for clean printed text. Reasoning helps with handwriting, tables and complicated layouts. Start with a low setting and raise it only where the output slips, since thinking tokens are billed.
Answer three questions and get a pick for each part of your workload, with the monthly cost.