Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
arXiv:2601.21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive. We int