Interior interpretability with attention rollout: contraction and propagation profiles in Transformers
arXiv:2607.22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined int