AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,450 results
27 Jul 2026

RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention

HardwareDGX agent

arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clus

25 Jul 2026

Can AMD break the CUDA Moat? AMD Advancing AI 2026

HardwareDGX agent

AMD’s AI accelerator software stack has progressed sharply, moving from a 0 % chance of catching Nvidia’s CUDA moat in early 2025 to a “great chance” of success by July 2026 after leadership changes a

24 Jul 2026

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
HardwareDGX agent

arXiv:2607.20940v1 Announce Type: new Abstract: Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoi

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

HardwareDGX agent

arXiv:2607.21553v1 Announce Type: new Abstract: We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate h

23 Jul 2026

AMD debuts next-generation AI infrastructure for frontier models, agentic workloads and autonomous robots

HardwareDGX agent

Advanced Micro Devices Inc. is pushing harder than ever to grab even more market share from Nvidia Corp. in the artificial intelligence chip industry. At its Advancing AI 2026 event today in San Franc

From Pixel to Prognosis: Convolutional and GLCM Feature Fusion for Automated Four-Class Cataract Severity Classification

HardwareDGX agent

arXiv:2607.18349v1 Announce Type: new Abstract: Objective: To develop a low-cost automated cataract severity classification system operating on standard consumer-grade colour photographs of the eye, w

Integrity of peer-to-peer distributed LLM inference under malicious nodes

ResearchDGX agent

arXiv:2607.19490v1 Announce Type: cross Abstract: Peer-to-peer distributed inference executes a Large Language Model (LLM) on pooled consumer hardware by spreading its layers across many nodes. Every

Leveraging ECRAM for Edge Continual Learning

HardwareDGX agent

arXiv:2607.19661v1 Announce Type: cross Abstract: Several edge computing platforms, such as autonomous vehicles and smart sensing devices, need to adapt to dynamic environments in real time by learnin

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

HardwareDGX agent

arXiv:2607.19456v1 Announce Type: cross Abstract: We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the fo

21 Jul 2026

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

HardwareDGX agent

The NVIDIA GB300 NVL72 platform set a world record by delivering 1,648 TFLOPs per GPU during the pre‑training of DeepSeek‑V3 671B, a mixture‑of‑experts (MoE) model. This achievement leveraged fifth‑ge

The first Vera Rubin clusters are here! Yesterday, @IneffableLabs took delivery of their Vera Rubin NVL72 cluster from @googlecloud @nvidia …

HardwareDGX agent

The first Vera Rubin clusters are here! Yesterday, @IneffableLabs took delivery of their Vera Rubin NVL72 cluster from @googlecloud @nvidia The AI frontier jumps forward by yet another generation of h

20 Jul 2026

NVIDIA NVLink: The Scale-Up Network for AI Factories

HardwareDGX agent

NVIDIA’s sixth‑generation NVLink is a purpose‑built scale‑up network that delivers up to 3.6 TB/s per GPU and 260 TB/s rack‑level bandwidth with 130 TFLOPS in‑network compute—outperforming off‑the‑she

16 Jul 2026

Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting

HardwareDGX agent

arXiv:2607.13808v1 Announce Type: new Abstract: Recent extensions of 3D Gaussian Splatting (3DGS) capture fine color details using hash-grid-based appearance parameterization but incur high computatio

9 Jul 2026

BREAKING: Dylan Patel (@dylan522p) of @SemiAnalysis_ says 'chips in Europe have less seasoning than chips in America & Mexico.' Plus: › Data…

HardwareDGX agent

BREAKING: Dylan Patel (@dylan522p) of @SemiAnalysis_ says 'chips in Europe have less seasoning than chips in America & Mexico.' Plus: › Data centers & France’s nuclear power › AI infrastructure over-o

InferNet: Exploiting Aggregate GPU Profiles as Side-Channel for DNN Architecture Inference

HardwareDGX agent

arXiv:2304.03388v2 Announce Type: replace Abstract: Deep Neural Networks (DNNs) have become ubiquitous for their ability to solve problems across various domains, including computer vision, natural la

Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization

HardwareDGX agent

arXiv:2607.06922v1 Announce Type: cross Abstract: Deep learning applications have been widely adopted on edge devices, to mitigate the privacy and latency issues of accessing cloud servers. Deciding t

8 Jul 2026

C4N, now GA: Delivering cloud’s highest per vCPU network and block storage I/O for x86 workloads

HardwareDGX agent

As organizations scale modern workloads — from high-throughput databases and network/security appliances to real-time analytics and AI/ML inference — network and block storage performance can quickly

7 Jul 2026

A Multi-Task Deep Learning Framework for Real-Time Intelligent Video Surveillance with Temporal Event Validation

HardwareDGX agent

arXiv:2607.03131v1 Announce Type: cross Abstract: Modern video surveillance systems generate far more video streams than human operators can effectively monitor, making automated analysis essential fo

Generative wave propagator

HardwareDGX agent

arXiv:2607.04440v1 Announce Type: cross Abstract: Seismic wavefield simulation is fundamental to seismology, but conventional finite-difference (FD) methods remain limited by numerical dispersion and

HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion

HardwareDGX agent

arXiv:2607.03012v1 Announce Type: cross Abstract: Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability to produce l

Scaling Weisfeiler-Leman Expressiveness Analysis to Massive Graphs with GPUs

HardwareDGX agent

arXiv:2607.02603v1 Announce Type: cross Abstract: The stable coloring of the Weisfeiler-Leman (1-WL) test is a cornerstone of Graph Neural Networks because it provides an upper bound to the expressive

The S-ICDF Dataset: Sionna-Simulated Dynamic Interference Characterization and Direction Finding

HardwareDGX agent

arXiv:2607.03411v1 Announce Type: cross Abstract: Jamming and spoofing threaten wireless and satellite navigation by disrupting or manipulating radio frequency (RF) signals, undermining availability,

Wan-Streamer v0.2: Higher Resolution, Same Latency

HardwareDGX agent

arXiv:2607.04443v1 Announce Type: cross Abstract: We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 mod

5 Jul 2026

183/365 of GPU Programming This 4.5 hour lesson on CUDA + ThunderKittens by @bfspector (TK co-author, Stanford PhD student) is one of the be…

HardwareDGX agent

183/365 of GPU Programming This 4.5 hour lesson on CUDA + ThunderKittens by @bfspector (TK co-author, Stanford PhD student) is one of the best educational videos on kernels out there (think @karpathy

3 Jul 2026

DeadPool: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint

HardwareDGX agent

arXiv:2607.01646v1 Announce Type: new Abstract: State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures acro

GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures

HardwareDGX agent

arXiv:2607.01409v1 Announce Type: cross Abstract: GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reconnecting ho

2 Jul 2026

“Within the next 18 months, you will be able to host GLM 5.2 equivalent intelligence on an RTX 5090 GPU.” -Ahmad Osman, AI World’s Fair

HardwareDGX agent

Ahmad Osman stated at AI World's Fair that within 18 months, GLM 5.2-equivalent AI intelligence will be deployable locally on a single RTX 5090 GPU, indicating rapid progress toward running advanced l

1 Jul 2026

AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance

HardwareDGX agent

arXiv:2606.30949v1 Announce Type: new Abstract: High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challen

Hierarchical Global Attention (HGA)

HardwareDGX agent

arXiv:2606.30709v1 Announce Type: cross Abstract: Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preserves the ori

27 Jun 2026

America is about to lose the AI race, and it will not happen at the frontier. It will happen at the floor. Everyone is watching who ships th…

HardwareDGX agent

America is about to lose the AI race, and it will not happen at the frontier. It will happen at the floor. Everyone is watching who ships the smartest model. The actual war is over the 80% of tokens n

26 Jun 2026

EGG: An Expert-Guided Agent Framework for Kernel Generation

HardwareDGX agent

arXiv:2606.26758v1 Announce Type: new Abstract: High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their developm

https://huggingface.co/nvidia/GLM-5.2-NVFP4

HardwareDGX agent

NVIDIA's GLM-5.2-NVFP4 is a quantized version of a large language model optimized for inference efficiency using NVIDIA's proprietary quantization format. The model is hosted on Hugging Face and repre

“If there is a deflation of the AI bubble, the optimists say that the new infrastructure will remain even if the companies do not — just as …

HardwareDGX agent

“If there is a deflation of the AI bubble, the optimists say that the new infrastructure will remain even if the companies do not — just as railways survived the 19th-century railway bust. However, th

SOLAR: AI-Powered Speed-of-Light Performance Analysis

ResearchDGX agent

arXiv:2606.26383v1 Announce Type: cross Abstract: How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to sof

24 Jun 2026

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

HardwareDGX agent

arXiv:2606.24369v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregr

End-to-End Radar and Communication Modulation Recognition with Neuromorphic Computing

HardwareDGX agent

arXiv:2606.24075v1 Announce Type: cross Abstract: Although deep learning-based methods can achieve high accuracy in automatic modulation recognition (AMR) tasks, their high computational cost makes it

OpenAI and Broadcom unveil LLM-optimized inference chip

HardwareDGX agent

OpenAI and Broadcom announced a collaboration on a specialized inference chip designed to optimize the performance and efficiency of large language model deployments. The chip, referred to as 'Jalapen

TurboMPC: Fast, Scalable, and Differentiable Model Predictive Control on the GPU

HardwareDGX agent

arXiv:2606.24039v1 Announce Type: new Abstract: Robotics increasingly relies on GPUs for parallel simulation, large-scale learning, and neural-network inference. For model predictive control (MPC) to

23 Jun 2026

CITADEL: CSI-Based Jamming Detection and Open-Set Classification for IIoT Networks

HardwareDGX agent

arXiv:2606.22939v1 Announce Type: cross Abstract: Radio frequency jamming poses a critical threat to the availability of wireless Industrial Internet of Things (IIoT) networks. Existing detection and

Fast-TurboQuant: A Multiplier-Free Online Vector Quantization Approach

HardwareDGX agent

arXiv:2606.21448v1 Announce Type: new Abstract: As large language models scale, memory bandwidth for key-value caches and retrieval-augmented generation systems becomes a critical bottleneck. While 1-

Real-World Deployment of Massively Parallel Sampling-Based MPC for Contact-Rich Manipulation

HardwareDGX agent

arXiv:2606.20712v1 Announce Type: new Abstract: Sampling-based Model Predictive Control (SMPC) is a promising strategy for contact-rich robotic manipulation, combining gradient-free optimization with

SCENIC: Semantic-Conditioned Edge-Aware Neural Framework for Structured IoT Command Generation

HardwareDGX agent

arXiv:2606.22296v1 Announce Type: new Abstract: Edge Internet of Things (IoT) agents are often constrained by memory capacity, privacy requirements, communication latency, and recurring inference cost

22 Jun 2026

The Secret Life of Computers

HardwareDGX agent

This article likely explores the hidden operational processes and internal workings of computers that users typically don't observe, such as microprocessor functions, system-level operations, or the b

11 Jun 2026

MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs

HardwareDGX agent

arXiv:2512.22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference int

9 Jun 2026

SwiftVR: Real-Time One-Step Generative Video Restoration

HardwareDGX agent

arXiv:2606.09516v1 Announce Type: new Abstract: Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusi

Toward Compiler World Models: Learning Latent Dynamics for Efficient Tensor Program Search

HardwareDGX agent

arXiv:2606.09312v1 Announce Type: new Abstract: Tensor program optimization is essential for modern machine learning systems, but its search space is enormous. Existing auto-schedulers reduce measurem

8 Jun 2026

E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory

HardwareDGX agent

arXiv:2601.16622v2 Announce Type: replace-cross Abstract: Equivariant Graph Neural Networks (EGNNs) have become a widely used approach for modeling 3D atomistic systems. However, mainstream architectu

6 Jun 2026

ITP-STDP: An Intrinsic-Timing Power-of-Two Learning Engine for On-Chip SNN Training

ResearchDGX agent

arXiv:2606.06159v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attention across

5 Jun 2026

Accelerating and Scaling MPC-Guided Reinforcement Learning for Humanoid Locomotion and Manipulation

HardwareDGX agent

arXiv:2606.05687v1 Announce Type: new Abstract: In humanoid motion control, model predictive control (MPC) offers physically grounded prediction and constraint handling, while reinforcement learning (

Staged Factorial Screening for Budget-Constrained Micro-Pretraining

HardwareDGX agent

arXiv:2606.05186v1 Announce Type: cross Abstract: Budget-constrained micro-pretraining often requires triaging many candidate recipes on a shared accelerator before larger search budgets are spent. We

Towards Realistic 3D Sonar Simulation

HardwareDGX agent

arXiv:2606.06130v1 Announce Type: new Abstract: As underwater robotics research increasingly addresses complex 3D perception and autonomous navigation, the fidelity of sonar simulation has become a ke

3 Jun 2026

CRAM-ER: Error-Resilient Spintronic Computational Random Access Memory for Scalable In-Memory Computation

HardwareDGX agent

arXiv:2606.02781v1 Announce Type: cross Abstract: Deep neural networks (DNNs) have achieved state-of-the-art performance across diverse domains. However, typical Von Neumann compute paradigms face sev

Releasing vui an open source voice mode 300M TTS model Runs on a single consumer gpu / apple sillicon Context aware speech 6 minutes of cont…

HardwareDGX agent

Jeremy Howard announced the release of Vui, an open-source voice mode text-to-speech (TTS) model with 300 million parameters that can run on consumer GPUs and Apple Silicon. The model features context

What’s new in serverless Managed Service for Apache Spark

HardwareDGX agent

Whether you use it for data preparation, real-time interactive queries, AI model training, or something entirely different, running Apache Spark at scale is demanding — you shouldn’t have to manage th

2 Jun 2026

FLARE: Diffusion for Hybrid Language Model

HardwareDGX agent

arXiv:2606.01774v1 Announce Type: cross Abstract: Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck for low-laten

Threshold-Based Exclusive Batching for LLM Inference

HardwareDGX agent

arXiv:2606.00516v1 Announce Type: new Abstract: Mixed batching (MB)--interleaving prefill and decode in a single batch--has become the standard scheduling strategy for large language model (LLM) infer

1 Jun 2026

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning

HardwareDGX agent

arXiv:2603.09221v2 Announce Type: replace Abstract: Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by projecting future states and selecting goal

On Efficient Scaling of GNNs via IO-Aware Layers Implementations

HardwareDGX agent

arXiv:2605.31500v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are bottlenecked by sparse, irregular memory access. Popular frameworks such as DGL and PyTorch Geometric support general

QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits

Model ReleasesDGX agent

arXiv:2605.30358v1 Announce Type: new Abstract: Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, where the performance is highly constrained to noise. Addressing the limit

Unbox one of NVIDIA's first co-packaged optics samples with us. See why we bet on CPO early.

HardwareDGX agent

When we design large GPU clusters, the network is no longer a background system. It's part of the compute envelope. At the 800G and NVIDIA GB300 NVL72 scale, the back-end fabric accounts for 86% of ne

← Previous
1…678910…75
Next →