MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
DGX agentarXiv:2605.26343v1 Announce Type: new Abstract: Mechanistic interpretability has identified small sets of attention heads that implement specific behaviours in transformer language models, but recover