Model Releases
There's a lot of real-world documents that are scanned, rotated, handwritten, or some combination of any of these elements. This week we cre…
There's a lot of real-world documents that are scanned, rotated, handwritten, or some combination of any of these elements. This week we created a comprehensive document extraction benchmark that cont
There's a lot of real-world documents that are scanned, rotated, handwritten, or some combination of any of these elements. This week we created a comprehensive document extraction benchmark that contains documents tagged with various "perception challenges", along with other tags denoting task challenges, table structure, business domain. These docs include regulatory filings, hand-filled tax forms, photocopied docs, sensor noise, and more. Codex is surprisingly good at scans, but not great on rotated docs. OCR solutions are reasonable on rotations/handwriting but struggle on more general scans. Check out ExtractBench! ArXiv: https://arxiv.org/pdf/2607.29677 Site: https://www.extractbench.ai/ Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we tested 14 systems on documents that weren't born digital: 1950s regulatory filings, hand-filled tax forms, and pages degraded with fax thresholding, photocopier tone curves, sensor …
Related
- We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑🔬 , our effort to create the most comprehensive, schema-guided, real-world document …
- ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…
- Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our app…
- Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models a…
Source: Jerry Liu (X) | 2026-08-14