Model Releases

Introducing ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—•๐—ฒ๐—ป๐—ฐ๐—ต: the most comprehensive benchmark for information extraction from complex enterprise documents. Our appโ€ฆ

Introducing ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—•๐—ฒ๐—ป๐—ฐ๐—ต: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems โ€” frontier VLMs, coding agents, ex

DGX agentx-post
model-releasesjerry-liu--x

Introducing ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—•๐—ฒ๐—ป๐—ฐ๐—ต: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems โ€” frontier VLMs, coding agents, extraction APIs โ€” on 370 enterprise docs, 4,869 pages, 67 doc types. Zero LLM judges, fully deterministic. Biggest finding: past 50 pages, commercial VLMs collapse below 35% recall. Precision stays high, but they silently drop most of the table rows. What is your extraction agent missing? Run ExtractBench to see today. Blog: https://www.llamaindex.ai/blog/introducing-extractbench GitHub: https://github.com/run-llama/ExtractBench HuggingFace: https://huggingface.co/datasets/llamaindex/ExtractBench Media

Related

Source: Jerry Liu (X) | 2026-08-11

Loading related sourcesโ€ฆ