Research
Deep Learning Models Also Recall Features
arXiv:2608.20970v1 Announce Type: new Abstract: Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual r
arXiv:2608.20970v1 Announce Type: new Abstract: Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call feature recall. The core observation is that a linear projection can be read as retrieving stored information scaled by input activations. I define feature recall, show it applies across architectures, and contrast it with the established paradigm of feature combination. I also consider how cases of feature recall might be mechanistically identified. The account gives philosophers a new conceptual tool for understanding deep learning, and points to empirical directions for mechanistic interpretability research.
Related
- Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
- Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
- LLM Self-Recognition: Steering and Retrieving Activation Signatures
- GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
Source: arXiv cs.AI | 2026-08-24