Model Releases

We benchmarked 20+ open-weight models on easy-to-hard document extraction tasks through ExtractBench. The results are all available on @hugg…

We benchmarked 20+ open-weight models on easy-to-hard document extraction tasks through ExtractBench. The results are all available on @huggingface 🤗 ExtractBench is a schema-guided extraction benchma

DGX agentx-post
model-releasesjerry-liu--x

We benchmarked 20+ open-weight models on easy-to-hard document extraction tasks through ExtractBench. The results are all available on @huggingface 🤗 ExtractBench is a schema-guided extraction benchmark that contains 4.8k+ pages across 8 domains and 67 document types, with a mix of short, medium, long docs and simple/complex schema.s The results reported is a measure of “value accuracy” through unified F1. ✅ Qwen 3.8 leads the pack ✅ earlier generations of Qwen models are also quite strong ✅ kimi-k3 is the next best. GLM-5.3-flash and qwen 3.8 flash also just released today - hopefully will have results on these soon! Note : these results don’t include visual grounding (whether each value is mapped to the right bounding box) and confidence scores. When you include these, it is increasingly clear why specialized OCR tools (like LlamaParse) matter, since this is metadata that’s hard to DIY by prompting the raw model. Come check out our HF leaderboard: https://huggingface.co/datasets/llamaindex/ExtractBench?leaderboard_base_model=false Learn more about ExtractBench here: https://www.extractbench.ai/

Related

Source: Jerry Liu (X) | 2026-08-26

Loading related sources…