CompanyOpenAI1 recent entries14 Jul 2026Post-Train NVIDIA Cosmos 3 in One Day Using Agent SkillsNVIDIA Cosmos 3 was post‑trained in under a day using TAO agent skills and LoRA adapters, raising accuracy on the Woven Traffic Safety video QA dataset from 54.41 % to 93.35 %. The mixture‑of‑transfor
CompanyGoogle1 recent entries9 Apr 2026How to Accelerate Protein Structure Prediction at Proteome-ScaleNVIDIA, Google DeepMind, EMBL-EBI, and Seoul National University collaborated to extend the AlphaFold Protein Structure Database (AFDB) beyond monomeric structures to proteome-scale quaternary stru...
CompanyMeta3 recent entries23 Jun 2026Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative DecodingDFlash is an open source block diffusion model for speculative decoding that significantly accelerates LLM inference on NVIDIA Blackwell GPUs by drafting entire token blocks in parallel and verifying →4 Aug 2026Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 SuperNVIDIA Alpamayo 2 Super is a publicly available 34‑billion‑parameter vision–language–action model that merges a 32‑B Cosmos 3 Super Reasoner with a 2‑B action‑expert diffusion network. It produces uni→10 Aug 2026Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIAMeta released Muse Glimmer, a 30‑billion‑parameter dense language model with a context window exceeding 120 K tokens, designed for local, long‑running agentic AI workloads. The model is optimized to r
CompanyDeepSeek3 recent entries24 Apr 2026Build with DeepSeek V4 Using NVIDIA Blackwell and GPU-Accelerated EndpointsDevelopers can build with DeepSeek V4 through NVIDIA GPU-accelerated endpoints on build.nvidia.com, with hosted endpoints providing a fast way to prototype before moving to self-hosted deployment. Dee→21 Jul 2026Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72The NVIDIA GB300 NVL72 platform set a world record by delivering 1,648 TFLOPs per GPU during the pre‑training of DeepSeek‑V3 671B, a mixture‑of‑experts (MoE) model. This achievement leveraged fifth‑ge→12 Aug 2026Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context
CompanyNVIDIA8 recent entries4 Aug 2026Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 SuperNVIDIA Alpamayo 2 Super is a publicly available 34‑billion‑parameter vision–language–action model that merges a 32‑B Cosmos 3 Super Reasoner with a 2‑B action‑expert diffusion network. It produces uni→4 Aug 2026Beyond VLAs: How World Action Models Reshape Robot ManipulationWorld Action Models (WAMs) use video-based world modeling instead of vision‑language backbones, giving robots a learned physics engine that supports zero‑shot transfer to new tasks, environments, and →10 Aug 2026Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIAMeta released Muse Glimmer, a 30‑billion‑parameter dense language model with a context window exceeding 120 K tokens, designed for local, long‑running agentic AI workloads. The model is optimized to r→11 Aug 2026Route AI Agent Workloads Across Models with NVIDIA NeMo SwitchyardNVIDIA NeMo Switchyard is a routing platform that directs AI agent workloads to the most suitable specialized or frontier model for each step of a task, balancing performance, cost, and latency. It of→11 Aug 2026NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running AgentsNVIDIA Nemotron 3.5 Lightning is an open‑source mixture‑of‑experts language model totaling 30 B parameters with only 3 B active during inference, designed to serve high‑volume, low‑latency execution f→11 Aug 2026NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 EmulationNVIDIA JetPack 7.2.1 adds agentic video skills with the unified jetson‑videosdk, allowing programmable, device-aware video workflows that link developer intent to live device discovery and performance→12 Aug 2026Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context →12 Aug 2026How to Choose Full-Stack Observability for NVIDIA AI FactoriesA full‑stack observability framework for NVIDIA AI factories links telemetry from compute, networking, storage, orchestration and application layers using specialized tools (DCGM, NVSM, UFM, NetQ, NMX