Model Releases
AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification
arXiv:2601.03605v2 Announce Type: replace Abstract: Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for mor
arXiv:2601.03605v2 Announce Type: replace Abstract: Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verification. Existing factuality verification methods do not capture graded judgments, even though factuality is better understood as a spectrum rather than a binary of right and wrong. To bridge this gap, we focus on graded factuality verification and propose AEScorer, an agentic evidence-grounded framework with two stages: agentic evidence acquisition and graded scoring. AEScorer first gathers and refines external evidence through agentic search, and then predicts a scalar factuality score to distinguish nuanced differences in factual correctness. We further construct GradedVeriBench, a benchmark for graded factuality verification spanning both general and multi-hop question answering. Experimental results on GradedVeriBench show that AEScorer substantially outperforms existing methods across both settings, demonstrating the value of coupling targeted evidence acquisition with graded scoring.
Related
- MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
- Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation
Source: arXiv cs.CL | 2026-08-28