Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
DGX agentarXiv:2601.12499v2 Announce Type: replace Abstract: Despite scaling to massive context windows, Large Language Models (LLMs) struggle with multi-hop reasoning due to inherent position bias, which caus