Model Releases

We wrote a 36-page ArXiv whitepaper on ExtractBench ๐Ÿง‘โ€๐Ÿ”ฌ , our effort to create the most comprehensive, schema-guided, real-world document โ€ฆ

We wrote a 36-page ArXiv whitepaper on ExtractBench ๐Ÿง‘โ€๐Ÿ”ฌ , our effort to create the most comprehensive, schema-guided, real-world document extraction benchmark. Itโ€™s extremely detailed and covers every

DGX agentx-post
model-releasesjerry-liu--x

We wrote a 36-page ArXiv whitepaper on ExtractBench ๐Ÿง‘โ€๐Ÿ”ฌ , our effort to create the most comprehensive, schema-guided, real-world document extraction benchmark. Itโ€™s extremely detailed and covers everything from comparisons with related work on document extraction, to the dataset construction / how ground-truth is generated, to our experiments over 14+ extraction systems. Here are some of the most salient points from the paper: โœ… The benchmark scores schema-guided extraction on real enterprise documents. Given input doc + schema and predicted output from an extractor, the benchmark measures value accuracy, grounding, tags for each โ€œchallengeโ€, and cost. โœ… Compared to other benchmarks, we have more schemas, more evaluation dimensions, and more data domain diversity โœ… The ground-truth is constructed according to 3 doc subtypes: real-docs use a model ensemble + human review, synthetic long lists have ground-truth by construction, and scanned forms also have human review incl. boxes โœ… Our three modes (LlamaParse cost-effective, agentic, and agentic plus) are at the Pareto frontier of accuracy and cost. The full ArXiv paper is here: https://arxiv.org/pdf/2607.29677 Our site: https://www.extractbench.ai/ Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A โ€ฆ

Related

Source: Jerry Liu (X) | 2026-08-12

Loading related sourcesโ€ฆ