Model Releases

ParseBench is here!📊 We’ve just released ParseBench, an open benchmark + dataset for evaluating document parsing at scale. It includes: • 2…

ParseBench is here!📊 We’ve just released ParseBench, an open benchmark + dataset for evaluating document parsing at scale. It includes: • 2,000+ human-reviewed enterprise documents • 167,000 evaluatio

DGX agentx-post
model-releasesjerry-liu--x

ParseBench is here!📊 We’ve just released ParseBench, an open benchmark + dataset for evaluating document parsing at scale. It includes: • 2,000+ human-reviewed enterprise documents • 167,000 evaluation rules • Coverage across 5 key areas: tables, charts, content faithfulness, semantic formatting, and visual grounding What makes it different? ParseBench optimizes for semantic correctness, not exact text matching. That means evaluating whether parsed outputs are actually useful for humans and AI agents making downstream decisions 🔍 Explore more: 📖 Blog: http://www.llamaindex.ai/blog/parsebench 💻 Code: http://github.com/run-llama/ParseBench 🤗 Dataset: http://huggingface.co/datasets/llamaindex/ParseBench

Source: Jerry Liu (X) | 2026-04-13

Loading related sources…