Agents

We comprehensively benchmarked Opus 4.8 on document understanding tasks, and compared it to Opus 4.7. It's fairly apparent that Opus 4.8 was…

We comprehensively benchmarked Opus 4.8 on document understanding tasks, and compared it to Opus 4.7. It's fairly apparent that Opus 4.8 wasn't explicitly post-trained on visual document understanding

DGX agentx-post
agentsjerry-liu--x

We comprehensively benchmarked Opus 4.8 on document understanding tasks, and compared it to Opus 4.7. It's fairly apparent that Opus 4.8 wasn't explicitly post-trained on visual document understanding: it does slightly better on tables/semantic formatting/layout, but worse on content faithfulness and more. Full results ready on ParseBench: https://www.parsebench.ai/ Opus 4.8 dropped today. ParseBench results are out. ✅ Slight gains: tables, semantic formatting, layout ⚠️ Slight regressions: charts, content faithfulness 💰 Slight price/page increase Lots of alpha left in teaching LLMs to read docs like humans do. LlamaParse remains the best d…

Source: Jerry Liu (X) | 2026-05-29

Loading related sources…