SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
arXiv:2606.08496v1 Announce Type: cross Abstract: Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse featur