OCR mini-bench

Business-first OCR benchmark for standard operational documents.

This benchmark compares OCR extraction performance on real business documents across repeated runs, so you can see both quality and consistency, not just a single score. It explicitly measures how well models transform an input document and expected keys into the correct output values. It highlights what matters in production: critical-field success, reliability over repeats (pass^n), latency, stability, and cost per successful outcome. Every model is run at its lowest thinking mode for a level playing field, unless noted otherwise (*).

42 docs • 31 models • 13,020 runs

Last updated: 22/07/2026

# Model Success pass^3 pass^5 Cost/success Latency Critical All fields Cost/doc Variance
1
Gemini 3 Flash
Google • BALANCED
73.8% 73.8% 73.8% 0.67¢ 16.0s 96.7% 95.3% 0.46¢
2
Claude Sonnet 4.6
Anthropic • BALANCED
73.8% 73.8% 73.8% 3.61¢ 18.9s 94.3% 94.7% 2.46¢
3
Gemini 3.6 Flash medium thinking
Google • BALANCED
73.8% 66.7% 62.9% 1.82¢ 28.0s 95.4% 92.7% 1.29¢
4
Claude Opus 4.6
Anthropic • SOTA
73.0% 72.2% 71.8% 6.25¢ 18.9s 95.8% 94.8% 4.15¢
5
GPT-5.6 Sol
OpenAI • SOTA
71.5% 64.3% 62.3% 8.28¢ 20.7s 96.3% 96.0% 5.97¢
6
Gemini 3.1 Pro
Google • SOTA
68.7% 68.7% 68.7% 2.55¢ 65.3s 91.5% 89.5% 1.63¢
7
Claude Opus 5
Anthropic • SOTA
68.6% 65.1% 64.4% 5.98¢ 17.4s 95.9% 94.7% 4.09¢
8
Claude Opus 4.7
Anthropic • SOTA
66.1% 63.0% 62.4% 6.59¢ 18.7s 95.2% 94.6% 4.13¢
9
Gemini 3.6 Flash
Google • BALANCED
65.9% 60.0% 58.4% 2.31¢ 13.7s 95.7% 93.8% 1.39¢
10
Claude Sonnet 5
Anthropic • BALANCED
64.6% 56.4% 53.8% 4.96¢ 18.2s 94.3% 93.1% 3.09¢
11
Gemini 3.5 Flash-Lite
Google • BUDGET
63.4% 53.9% 49.7% 0.69¢ 10.6s 95.1% 92.7% 0.42¢
12
GPT-5.6 Terra
OpenAI • SOTA
62.8% 55.3% 52.1% 4.95¢ 14.6s 95.5% 95.0% 3.00¢
13
GPT-5.6 Luna
OpenAI • BALANCED
62.8% 52.9% 49.9% 1.90¢ 13.6s 95.0% 94.3% 1.20¢
14
GPT-5.6 Luna medium thinking
OpenAI • BALANCED
61.9% 52.4% 48.7% 2.16¢ 15.0s 94.9% 94.3% 1.34¢
15
Gemini 3.1 Flash-Lite
Google • BALANCED
61.2% 61.2% 61.2% 0.32¢ 12.8s 93.3% 93.2% 0.19¢
16
GPT-5.6 Terra medium thinking
OpenAI • SOTA
60.4% 54.6% 51.2% 5.23¢ 13.4s 95.0% 94.8% 3.07¢
17
Gemini 2.5 Flash-Lite
Google • BUDGET
58.6% 58.6% 58.6% 0.10¢ 14.4s 94.1% 91.9% 0.06¢
18
GPT-5.5
OpenAI • SOTA
57.3% 47.3% 43.2% 9.23¢ 16.2s 94.2% 94.0% 4.13¢
19
Gemini 3.5 Flash
Google • BALANCED
57.3% 57.3% 57.3% 3.36¢ 13.8s 94.8% 93.9% 1.67¢
20
Medium
Mistral • BALANCED
54.1% 49.8% 47.5% 0.69¢ 21.0s 91.0% 88.4% 0.29¢
21
Gemini 3.5 Flash-Lite high thinking
Google • BUDGET
52.9% 43.5% 40.2% 0.84¢ 20.3s 93.5% 91.6% 0.42¢
22
Large
Mistral • SOTA
50.5% 48.4% 47.3% 0.31¢ 23.2s 92.0% 89.9% 0.28¢
23
GPT-5.4
OpenAI • SOTA
49.2% 38.8% 35.9% 7.82¢ 26.9s 92.7% 93.2% 2.22¢
24
OCR
Mistral • SOTA
48.4% 43.5% 41.6% 0.67¢ 11.8s 92.3% 91.8% 0.30¢
25
Small
Mistral • BUDGET
46.2% 43.1% 41.9% 0.12¢ 12.6s 88.6% 88.0% 0.05¢
26
GPT-5
OpenAI • SOTA
44.6% 39.3% 37.9% 24.20¢ 19.8s 88.8% 89.4% 1.01¢
27
GPT-5.4 mini
OpenAI • BALANCED
43.2% 35.9% 32.4% 3.30¢ 13.7s 91.5% 92.0% 0.65¢
28
GPT-5 mini
OpenAI • BALANCED
39.3% 32.7% 30.5% 3.09¢ 25.0s 90.2% 89.9% 0.28¢
29
Claude Haiku 4.5
Anthropic • BUDGET
34.9% 34.9% 34.9% 3.73¢ 13.6s 89.9% 89.9% 0.97¢
30
GPT-5.4 nano
OpenAI • BUDGET
23.6% 13.7% 11.3% 7.05¢ 19.9s 82.8% 78.7% 0.27¢
31
GPT-5 nano
OpenAI • BUDGET
8.7% 5.2% 4.1% 2.88¢ 17.0s 63.1% 52.6% 0.05¢

passn metric: Probability of n consecutive successes in n runs (strict).

Variance column: Shows min–max interval with bar width indicating spread.

All metrics: Aggregated across all documents and 10 runs per model.

(*) Thinking mode. Every model is run at its lowest thinking mode — this establishes the lower boundary of capability and keeps a level playing field, since default thinking budgets differ across models and providers. So: lowest thinking, unless noted otherwise. From 22/07 we also started testing models at their default / higher thinking modes to see how they compare against their non-thinking counterparts; those runs are tagged with a thinking badge and can be shown with the Thinking variants toggle above the table.