Model Releases
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
arXiv:2509.22750v3 Announce Type: replace Abstract: Real-world multi-hop QA is naturally linked with ambiguity, where a single query can trigger multiple reasoning paths that require independent resol
arXiv:2509.22750v3 Announce Type: replace Abstract: Real-world multi-hop QA is naturally linked with ambiguity, where a single query can trigger multiple reasoning paths that require independent resolution. Since ambiguity can occur at any stage, models must navigate layered uncertainty throughout the entire reasoning chain. Despite its prevalence in real-world user queries, previous benchmarks have primarily focused on single-hop ambiguity, leaving the complex interaction between multi-step inference and layered ambiguity underexplored. In this paper, we introduce extbf{MARCH}, a benchmark for their intersection, with 2,209 multi-hop ambiguous questions curated via multi-LLM verification and validated by human annotation with strong agreement. Our experiments reveal that even state-of-the-art models struggle with MARCH, confirming that combining ambiguity resolution with multi-step reasoning is a significant challenge. To address this, we propose extbf{CLARION}, a two-stage agentic framework that explicitly decouples ambiguity planning from evidence-driven reasoning, significantly outperforms existing approaches, and paves the way for robust reasoning systems.
Related
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
- Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
- FinTruthQA: A Benchmark for AI-Driven Financial Disclosure Quality Assessment in Investor -- Firm Interactions
- PeReGrINE: Evaluating Personalized Review Fidelity with User Item Graph Context
Source: arXiv cs.CL | 2026-04-10