From Mechanistic to Compositional Interpretability
DGX agentarXiv:2605.08934v1 Announce Type: new Abstract: Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable com