Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications
DGX agentarXiv:2603.22473v2 Announce Type: replace-cross Abstract: Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the