Safety

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders

arXiv:2604.03919v2 Announce Type: replace-cross Abstract: We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpretable

DGX agentpaper
safetyarxiv-cs-ai

arXiv:2604.03919v2 Announce Type: replace-cross Abstract: We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpretable, monosemantic features but destroy temporal coherence: hard TopK selection produces unstable feature assignments across frames, reducing autocorrelation by 36%. We propose spatio-temporal contrastive objectives and Matryoshka hierarchical grouping that recover and even exceed raw temporal coherence. The contrastive loss weight controls a tunable trade-off between reconstruction and temporal coherence. A systematic ablation on two backbones and two datasets shows that different configurations excel at different goals: reconstruction fidelity, temporal coherence, action discrimination, or interpretability. Contrastive SAE features improve action classification by +3.9 pp over raw features and text-video retrieval by up to 2.8 x R@1. A cross-backbone analysis reveals that standard monosemanticity metrics contain a backbone-alignment artifact: both DINOv2 and VideoMAE produce equally monosemantic features under an independent (CLIP) similarity space. Targeted feature ablation shows that contrastive training concentrates the probe's predictive signal into a small number of identifiable features. Supplementary material, code, configurations and evaluation scripts are available at https://github.com/atahandokme/spatio-temporal-sparse-autoencoders-video.

Source: arXiv cs.AI | 2026-08-11

Loading related sources…