Model Releases

NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents

NVIDIA Nemotron 3.5 Lightning is an open‑source mixture‑of‑experts language model totaling 30 B parameters with only 3 B active during inference, designed to serve high‑volume, low‑latency execution f

DGX agentarticle
model-releasesnvidia-developer

NVIDIA Nemotron 3.5 Lightning is an open‑source mixture‑of‑experts language model totaling 30 B parameters with only 3 B active during inference, designed to serve high‑volume, low‑latency execution for always‑on AI agents. It delivers up to four times the output speed of similarly sized models while preserving accuracy through speculative decoding, harness‑optimized training and NVFP4/BF16 quantization. The model integrates with NeMo Switchyard for intelligent routing across local or data‑center deployments and is released under a permissive license with full weights and recipes for community customization.

Related

Source: NVIDIA Developer | 2026-08-11

Loading related sources…