Hierarchical Global Attention (HGA)
arXiv:2606.30709v1 Announce Type: cross Abstract: Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preserves the ori
Knowledge catalogue
arXiv:2606.30709v1 Announce Type: cross Abstract: Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preserves the ori
Databricks implements reliability measures and monitoring systems to ensure consistent GPU performance and availability across its AI platform infrastructure. The article likely covers their approache
Lambda Labs presented a keynote address at the ALVR (Augmented Language and Vision Research) workshop, which was held in conjunction with ACL 2026, a major computational linguistics conference. The pr
arXiv:2606.31050v1 Announce Type: cross Abstract: How to accurately predict a high-fidelity future world? While the visual world is inherently continuous, existing deterministic video prediction model
Dylan Patel, a notable tech industry figure, was visited by the post author (Sassine Ghazi), and they had a positive conversation together. The post appears to be a casual social media update sharing
NVIDIA and its partners are investing in American manufacturing, supply chains, energy grids and skilled workforces so the U.S. can produce the infrastructure needed for better healthcare, breakthroug
Our research team has 9 papers at ICML next week! Spanning the full stack from frontier agents to GPU kernels, we're excited to share what our researchers and collaborators have been working on. If yo
Artificial intelligence chipmaking startup Oxmiq Labs Inc. says it wants to become the next Arm Holdings Plc. after raising 35 million in an early-stage A funding, bringing its total amount raised to
arXiv:2606.31145v1 Announce Type: new Abstract: Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with seq
Bloomberg: Taiwanese authorities detain two Super Micro staff and an Albatron manager after a raid of Super Micro's local offices this week over Nvidia shipments to China — Taiwanese prosecutors detai
Together AI Inc., the operator of a cloud platform optimized to run open-source artificial intelligence models, has raised 800 million from investors. The startup stated in its funding announcement to
Physical security company Verkada Inc. has taken an investment from Nvidia Corp. and signed a technical partnership with the chipmaker, the two said today, in a deal meant to speed up the artificial i
When boomer companies get high Anthropic bill, they set spend limits When I see a nearly million dollar Anthropic monthly bill, my first reaction is to complain about use of shitty Haiku models Imagin
arXiv:2606.28519v1 Announce Type: new Abstract: Training operator-learning models for large-scale problems governed by partial differential equations (PDEs) is challenging due to the curse of dimensio
Dina Bass / Bloomberg: AI inference startup Etched raised 800M from investors including Jane Street and a TSMC-linked venture firm, and says it has signed sales contracts worth 1B — Nvidia rival says
Apple. Anthropic. Disney Research. Google. Meta. Microsoft. NVIDIA. OpenAI. Few places outside Silicon Valley can claim R&D hubs from all of these companies. Fewer still are concentrated in a city of
arXiv:2606.28516v1 Announce Type: new Abstract: We present CLEAR-MoE, a four-phase post-training pipeline that converts a frozen pretrained Vision Transformer (ViT) into a sparse Mixture-of-Experts (M
arXiv:2606.28361v1 Announce Type: cross Abstract: Multi-step retrieval-augmented generation (RAG) has been widely deployed as LLM-powered web services for complex question answering, where iterative r
arXiv:2606.30293v1 Announce Type: new Abstract: Robotic applications increasingly rely on distributed computational infrastructures that combine embedded devices, edge servers, and cloud resources. Th
arXiv:2606.30292v1 Announce Type: cross Abstract: We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulation. The system adapts the LongLive
arXiv:2602.16634v2 Announce Type: replace-cross Abstract: The rare-event sampling problem has long been the central limiting factor in molecular dynamics (MD), especially in biomolecular simulation. R
arXiv:2606.29082v1 Announce Type: new Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrate
Featuring the man of the moment @dylan522p @shaunmmaguire in today's Training Data episode. Nobody is a more trusted industry insider to the biggest infrastructure build-out in history. The story of h
arXiv:2606.28394v1 Announce Type: new Abstract: The physical anastylosis of collapsed architectural monuments -- the meticulous reassembly of fallen stone elements into their original structural confi
arXiv:2606.30497v1 Announce Type: cross Abstract: We present a comparative study of CUDA optimization strategies applied to forward and backward propagation in a shallow neural network. Three stacked
arXiv:2606.28831v1 Announce Type: cross Abstract: Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top-p nucleus sampling) offer superior accuracy b
When Jaiveer Singh talks about robots, he doesn’t begin with spectacle. He begins with infrastructure: the boards inside machines, the software that lets developers see through a robot’s cameras and t
As organizations move from AI pilots to production AI factories, infrastructure decisions have shifted from peak chip specifications to cost per token: how many useful tokens they can deliver per doll
In this post, we explore how Outpost VFX achieved 8x faster training speeds using AWS infrastructure to transform their face replacement workflow, the technical architecture they implemented to overco
Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners, and enterprises can transform their workflows using the latest advances in OpenUSD and NVI
It's wild how quickly Etched designed and got the chips out, all within 2 years. They went deep, hardcoding attention into silicon and getting very high MFU. This kind of hardware tailored made for LL
arXiv:2606.29119v1 Announce Type: cross Abstract: We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle outer loop over
Amazon Web Services Inc.‘s next-generation silicon is working its way deeper into the company’s cloud compute infrastructure offerings with the debut of the new Amazon EC2 C9g and EC2 C9gd instances t
Agentic AI is driving key technology providers to rethink the computing architecture required to run rapidly expanding autonomous systems. In response to this challenge, two leading tech companies hav
Project gallery and whitepaper: https://research.nvidia.com/labs/gear/aspire/ ASPIRE is a great collaboration between NVIDIA GEAR lab, UMich, Berkeley, and CMU. Kudos to all the coauthors who pour the
arXiv:2606.29329v1 Announce Type: new Abstract: We study the problem of physically plausible shadow casting when animating 3D Gaussian Splatting (3DGS) avatars, either individually or in multi-avatar
arXiv:2606.30436v1 Announce Type: new Abstract: Scaling monocular 3D Gaussian Splatting (3DGS) SLAM to kilometer-level outdoor environments poses two tightly coupled challenges: fragile long-term pose
arXiv:2602.15727v2 Announce Type: replace-cross Abstract: Visual analogy learning enables image editing via demonstration rather than textual description, allowing users to specify complex transformat
arXiv:2403.07711v5 Announce Type: replace-cross Abstract: Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in extend
Chip component supplier Stathera Inc. today announced that it has closed a 35 million funding round. Semiconductor-focused fund Maverick Silicon led the Series A deal. It was joined by chipmaker Media
arXiv:2606.21401v2 Announce Type: replace-cross Abstract: Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their infer
TokenBudgeting examines enterprise spending patterns and cost management strategies related to AI token consumption, based on SemiAnalysis's direct conversations with corporate customers. The analysis
arXiv:2602.04940v2 Announce Type: replace Abstract: Deep learning has emerged as a transformative tool for the neural surrogate modeling of partial differential equations (PDEs), known as neural PDE s
arXiv:2601.16956v1 Announce Type: cross Abstract: The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has necessitate
arXiv:2606.27612v1 Announce Type: cross Abstract: Silicon photonics enables integration of optical components using standard semiconductor processes, greatly improving data communication bandwidth and
Dylan Patel discusses concerns within the AI startup ecosystem about potential Chinese AI model bans and their impact on companies generating significant revenue. The post appears to address industry-
Firefly Aerospace partnered with NVIDIA to deploy an NVIDIA Jetson module on its Elytra spacecraft in lunar orbit, enabling on-orbit AI processing for the Ocula Moon imaging service. The Jetson module
Going to be in Paris for @RaiseSummit? Join us and @nvidia to hear from leaders at both companies on the future of AI inference and infrastructure. After that: cocktails and a DJ. Oh la la! See you th
Enterprise autonomous AI agents inspect code, run tests, search documents, and operate for extended periods on behalf of users , requiring governance frameworks to control access and actions. The Ente
This newsletter issue covers three major topics in AI development: advances in self-improving robotic systems, details about a large-scale Chinese GPU cluster with 10,000 units for AI training, and a
Nvidia CEO Jensen “@Tesla stack is the most advanced autonomous vehicle stack in the world. I’m fairly certain they were already using end-to-end AI. Whether their AI did reasoning or not in somewhat
One day left to submit your projects for the Hermes Agent Accelerated Business Hackathon presented by @NVIDIAAI × @stripe × @NousResearch! Submissions close 11:59 PM PT tomorrow, June 30th. Last minut
Debby Wu / Bloomberg: Source: Taiwanese authorities raided Super Micro's Taiwan office as part of a probe into the alleged smuggling of Nvidia chips to China; SMCI closed down 8.1% — Super Micro Compu
arXiv:2606.27396v1 Announce Type: cross Abstract: Test-input generation for tensor kernels is folkloric. Most projects pick a representative shape and dtype, run a fixed-shape allclose-style check, an
arXiv:2606.27538v1 Announce Type: cross Abstract: We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes
Many organizations today are floating in a ton of data. This is good for powering AI agents, but less favorable when it comes to identifying which signals from all of that data actually matter. The so
Bloomberg: Australia-based Firmus partners with Nvidia to build its first data center in Batam, Indonesia; the 360 MW Nvidia DSX AI factory campus is developed with DayOne — Firmus Technologies Pty Lt
This post humorously contrasts the competitive achievements people in San Francisco boast about—such as winning the International Mathematical Olympiad, achieving high competitive programming ratings,
Don Clark / New York Times: A look at advanced chip packaging, now more reliant on TSMC and its partners in Taiwan than ever, and the efforts to address this bottleneck in the US — A silicon wafer ref
Politico: AI executives and lobbyists say they are seeking regulatory clarity from the Trump administration but are wary of pressing for answers, fearing retaliation — The unpredictability was on disp