Research
Real-time multilingual ASR using rolling buffers and monolingual models [P]
This research discusses a technique for enabling real-time multilingual automatic speech recognition (ASR) on edge devices by dynamically switching between compact monolingual models rather than using
This research discusses a technique for enabling real-time multilingual automatic speech recognition (ASR) on edge devices by dynamically switching between compact monolingual models rather than using a single large multilingual model. The approach reduces computational and memory requirements while preserving the recognition performance of monolingual ASR models and enabling multilingual interactions on edge devices. The system incorporates components like an Encoder Endpointer and End-of-Utterance Joint Layer to optimize quality and latency tradeoffs, supports code-switching in real time, and can run on mobile devices in less than real time.
Source: r/MachineLearning | 2026-06-01