Model Releases
Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA
Meta released Muse Glimmer, a 30‑billion‑parameter dense language model with a context window exceeding 120 K tokens, designed for local, long‑running agentic AI workloads. The model is optimized to r
Meta released Muse Glimmer, a 30‑billion‑parameter dense language model with a context window exceeding 120 K tokens, designed for local, long‑running agentic AI workloads. The model is optimized to run on NVIDIA devices—GeForce RTX 5090, DGX Spark, DGX Station, and Jetson—and delivers high reliability, long‑context coherence, and predictable latency without mixture‑of‑experts routing overhead, achieving around 20 k tokens/sec per GPU. Deployment can be managed via NVIDIA NIM containers, SGLang, or vLLM for efficient on‑device inference in tasks such as software automation and autonomous agents.
Related
- Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong pe…
- Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows
Source: NVIDIA Developer | 2026-08-10