Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
DGX agentarXiv:2601.08258v3 Announce Type: replace Abstract: Large language models increasingly fail in a way that scalar accuracy cannot diagnose: they produce a sound reasoning trace and then abandon it unde