Research
An Exploration of Mamba for Speech Self-Supervised Models
arXiv:2506.12606v2 Announce Type: replace Abstract: While Mamba has demonstrated strong performance in language modeling, its potential as a speech self-supervised learning (SSL) model remains underex
arXiv:2506.12606v2 Announce Type: replace Abstract: While Mamba has demonstrated strong performance in language modeling, its potential as a speech self-supervised learning (SSL) model remains underexplored, with prior studies limited to isolated tasks. To address this, we explore Mamba-based HuBERT models as alternatives to Transformer-based SSL architectures. Leveraging the linear-time Selective State Space, these models enable fine-tuning on long-context ASR with significantly lower compute. Moreover, they show superior performance when fine-tuned for streaming ASR. Beyond fine-tuning, these models show competitive performance on SUPERB probing benchmarks, particularly in causal settings. Our analysis shows that they yield higher-quality quantized representations and capture speaker-related features more distinctly than Transformer-based models. These findings highlight Mamba-based SSL as a promising and complementary direction for long-sequence modeling, real-time speech modeling, and speech unit extraction. The codebase is available at https://github.com/hckuo145/Mamba-based-HuBERT.
Related
- [[bd-tp-self-supervised-speech-models-discover-phonological-ve|[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic]]
- emg2speech: Synthesizing speech from electromyography using self-supervised speech models
- Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yoruba
Source: arXiv cs.CL | 2026-04-21