Model Releases
The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released Extract…
The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released ExtractBench yesterday: 370 enterprise docs, 14 systems. The hardes
The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released ExtractBench yesterday: 370 enterprise docs, 14 systems. The hardest test: long-list completeness. An unclaimed-property list with 26,725 rows. A creditor matrix with 8,624 records. A 13F with 3,063 holdings. Frontier VLMs don't misread these docs, they abandon them. Precision stays high, recall collapses: 8.9–35.8% F1 on the longest documents. Every row they return looks correct, so spot checks pass while most of the document never came back. Our new Extract tier, Agentic Plus, processes long docs iteratively instead of one pass: 96.1% F1 on long-list tasks, and the only system that holds flat as docs get longer. Learn more about ExtractBench below👇 Blog: https://www.llamaindex.ai/blog/introducing-extractbench Paper: https://arxiv.org/pdf/2607.29677 Media
Related
- ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…
- Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our app…
- We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑🔬 , our effort to create the most comprehensive, schema-guided, real-world document …
- Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models a…
Source: Jerry Liu (X) | 2026-08-12