Safety from Honesty in a Disinterested AI Predictor
DGX agentarXiv:2606.29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior th