Model Releases
ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…
ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document types, spanning 8 real-world domains: finance, energy, gov, auto
ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document types, spanning 8 real-world domains: finance, energy, gov, auto, supply chain, healthcare, legal, real estate ✅ It covers a distribution of short, medium, and long documents ✅ It covers a variety of very complex table edge cases: tables with over 1k rows, nested tables within cells, cross-page tables, and more ✅ It covers scans, handwriting, and rotated pages We benchmarked across 14 different VLMs, coding agents, and document extraction APIs. Everything is fully public on our blog, ArXiv, Github, and HuggingFace: Blog: https://www.llamaindex.ai/blog/introducing-extractbench Hugging Face: https://huggingface.co/datasets/llamaindex/ExtractBench Github: https://github.com/run-llama/ExtractBench ArXiv: https://arxiv.org/pdf/2607.29677 Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A …
Related
- Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models a…
- Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our app…
Source: Jerry Liu (X) | 2026-08-11