Model Releases
Introducing ๐๐ ๐๐ฟ๐ฎ๐ฐ๐๐๐ฒ๐ป๐ฐ๐ต: the most comprehensive benchmark for information extraction from complex enterprise documents. Our appโฆ
Introducing ๐๐ ๐๐ฟ๐ฎ๐ฐ๐๐๐ฒ๐ป๐ฐ๐ต: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems โ frontier VLMs, coding agents, ex
Introducing ๐๐ ๐๐ฟ๐ฎ๐ฐ๐๐๐ฒ๐ป๐ฐ๐ต: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems โ frontier VLMs, coding agents, extraction APIs โ on 370 enterprise docs, 4,869 pages, 67 doc types. Zero LLM judges, fully deterministic. Biggest finding: past 50 pages, commercial VLMs collapse below 35% recall. Precision stays high, but they silently drop most of the table rows. What is your extraction agent missing? Run ExtractBench to see today. Blog: https://www.llamaindex.ai/blog/introducing-extractbench GitHub: https://github.com/run-llama/ExtractBench HuggingFace: https://huggingface.co/datasets/llamaindex/ExtractBench Media
Related
- We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. โ It contains 2k pโฆ
- ParseBench is the most comprehensive OCR benchmark for real-world enterprise documents: financial filings, contracts, insurance documents, aโฆ
- ParseBench is the first benchmark to include VLM chart understanding ๐๐๐ over enterprise documents. ๐ Existing benchmarks (ChartQA, Charโฆ
- ParseBench is now live on @Kaggle. The first document OCR benchmark built for AI agents โ 2,000 enterprise pages, 167K+ test rules, 5 dimensโฆ
Source: Jerry Liu (X) | 2026-08-11