Model Releases
We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few point…
We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing
We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing charts and visual grounding. For reference, Gemini 3.6 Flash has better results on tables, is slightly worse on charts, and is half the price. LlamaParse agentic is better on all fronts (including tables) at 1/6th of the price. tl;dr use Opus 5 as much as you want for coding and knowledge work. But at 8c per page and average OCR performance, don't use it to parse documents at scale Our full ParseBench leaderboard which we periodically update is here: https://www.parsebench.ai/
Related
- We comprehensively benchmarked Opus 4.7 on document understanding. We evaluated it through ParseBench - our comprehensive OCR benchmark for …
- We need more evals for document understanding. ParseBench is a really great start. I respect @llama_index ‘s work on this. 📈
- We benchmarked Mistral OCR against other frontier and open-weight models on ParseBench 📊 For a model at its price point, it is quite compet…
- Anthropic says Opus 4.7 hits 80.6% on Document Reasoning — up from 57.1%. But 'reasoning about documents' ≠ 'parsing documents for agents.' …
- Thanks for the shoutout re: ParseBench! 📑 🙂 Opus 4.7 is definitely a step up from Opus 4.6 on document understanding capabilities. Great t…
Source: Jerry Liu (X) | 2026-07-24