The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
DGX agentarXiv:2607.12963v1 Announce Type: new Abstract: As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by lo