LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
DGX agentarXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the