Local Ai
ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face
GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Expe
GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding speech understanding. Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool-use, and text-only tasks. The temporal grounding skills are trained on TimeGround-1M — a purpose-built dataset of long-form audio paired with time-aligned annotations. arXiv : https://arxiv.org/abs/2607.10387 Full Paper : https://arxiv.org/pdf/2607.10387.pdf HF Dataset : https://huggingface.co/datasets/ai-sage/TimeGround-1M HF Space : https://huggingface.co/spaces/hugging-apps/gigachat-audio-10b-a1-8b-demo submitted by /u/pmttyji [link] [comments]
Related
- MiMo-V2.5-GGUF (preview available)
- FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence
- VoxCPM TTS model + LoRa training abilities right in Comfy
Source: r/LocalLLaMA | 2026-07-26