Disillusionment with mechanistic interpretability research [D]
DGX agentMechanistic interpretability research aims to uncover specific neurons and circuits in neural networks responsible for tasks, but over a decade of efforts suggests these findings may not translate int