Model Releases
ParseBench is here!📊 We’ve just released ParseBench, an open benchmark + dataset for evaluating document parsing at scale. It includes: • 2…
ParseBench is here!📊 We’ve just released ParseBench, an open benchmark + dataset for evaluating document parsing at scale. It includes: • 2,000+ human-reviewed enterprise documents • 167,000 evaluatio
ParseBench is here!📊 We’ve just released ParseBench, an open benchmark + dataset for evaluating document parsing at scale. It includes: • 2,000+ human-reviewed enterprise documents • 167,000 evaluation rules • Coverage across 5 key areas: tables, charts, content faithfulness, semantic formatting, and visual grounding What makes it different? ParseBench optimizes for semantic correctness, not exact text matching. That means evaluating whether parsed outputs are actually useful for humans and AI agents making downstream decisions 🔍 Explore more: 📖 Blog: http://www.llamaindex.ai/blog/parsebench 💻 Code: http://github.com/run-llama/ParseBench 🤗 Dataset: http://huggingface.co/datasets/llamaindex/ParseBench
Source: Jerry Liu (X) | 2026-04-13