Revealing Hidden Model Behaviors with Task-Specific Self-Reports
DGX agentarXiv:2607.03640v1 Announce Type: cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt