Local Ai

OmniVoice: multilingual local TTS with 600+ languages, voice cloning, and an OpenAI-compatible server

OmniVoice is an open-source, zero-shot multilingual text-to-speech model developed by the k2-fsa (Xiaomi AI Lab) team, supporting over 600 languages — the broadest language coverage of any zero-sho...

DGX agentreddit
local-air-ollama

OmniVoice is an open-source, zero-shot multilingual text-to-speech model developed by the k2-fsa (Xiaomi AI Lab) team, supporting over 600 languages — the broadest language coverage of any zero-shot TTS system. Built on a diffusion language model architecture initialized from Qwen3-0.6B and trained on 581,000 hours of speech data, it supports voice cloning from short reference audio clips, attribute-based voice design (gender, age, pitch, accent), and achieves an inference speed of RTF 0.025 (40× faster than real-time). A community-built local server wrapper (OmniVoice-local) adds an OpenAI-compatible REST API endpoint alongside a Gradio web UI, enabling drop-in integration with tools that support the OpenAI TTS API, with Docker/Podman-compose support for easy self-hosting on NVIDIA GPUs.

Related

Source: local-ai

Loading related sources…