Model Releases
NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents
NVIDIA Nemotron 3.5 Lightning is an open‑source mixture‑of‑experts language model totaling 30 B parameters with only 3 B active during inference, designed to serve high‑volume, low‑latency execution f
NVIDIA Nemotron 3.5 Lightning is an open‑source mixture‑of‑experts language model totaling 30 B parameters with only 3 B active during inference, designed to serve high‑volume, low‑latency execution for always‑on AI agents. It delivers up to four times the output speed of similarly sized models while preserving accuracy through speculative decoding, harness‑optimized training and NVFP4/BF16 quantization. The model integrates with NeMo Switchyard for intelligent routing across local or data‑center deployments and is released under a permissive license with full weights and recipes for community customization.
Related
- v0.32.9
- NVIDIA Nemotron 3.5 Lightning is now live on Together AI. The fastest open model in its class is built for always-on agents that need to com…
- Nemotron 3.5 Lightning is available in LM Studio! The model is 30B MoE (3B active), can run very fast, and is trained for high volume agenti…
Source: NVIDIA Developer | 2026-08-11