From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach
DGX agentarXiv:2605.21303v1 Announce Type: new Abstract: Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimen