Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
DGX agentarXiv:2608.09900v1 Announce Type: new Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably wa