Model Releases

Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we tested 14 systems on documents that weren…

Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we tested 14 systems on documents that weren't born digital: 1950s regulatory filings, hand-filled tax f

DGX agentx-post
model-releasesjerry-liu--x

Every document extraction system has a perception blind spot. We mapped them. For ExtractBench, we tested 14 systems on documents that weren't born digital: 1950s regulatory filings, hand-filled tax forms, and pages degraded with fax thresholding, photocopier tone curves, sensor noise, and phone-camera capture. The failures don't overlap. 💡 Codex reads scans and handwriting above 93%, then drops to ~80% on rotated or image-only pages. 💡 Specialized APIs are the exact inverse: fine on rotation and handwriting, 81% on scans 💡 Gemini 3.5 Flash falls from 88.6% to 71.1% the moment a page is scanned. You benchmark on clean PDFs. Production sends you a shadowed photocopy from 1953. Our new Extract Tier, Agentic Plus, was the only system with no blind spot: 95.9% / 93.9% / 93.8% across rotated, scanned, and handwritten, a 2-point spread where others swing 10+. Learn more about ExtractBench 👇️ Blog: https://lnkd.in/gNm97fXp Paper: https://lnkd.in/euAfScWx Media

Related

Source: Jerry Liu (X) | 2026-08-13

Loading related sources…