Research
DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such a
Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as “Which actor from the film Heat won at least one Academy Award?”, which requires (1) distinguishing between multiple films sharing the same title and (2) reasoning across a large set of actors to gather and integrate evidence. Existing QA benchmarks rarely evaluate both challenges jointly. To address this, we introduce DEEPAMBIGQAGEN, an automatic data generation pipeline that constructs QA…
Related
- Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering
- AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
Source: Apple ML Research | 2026-08-06