Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance
DGX agentarXiv:2605.31304v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are widely used, but interpreting what they actually learn remains difficult. A major obstacle is that individual neurons