FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
arXiv:2601.06199v3 Announce Type: replace-cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Unlike images or