AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,732 results
12 May 2026

Cluster magicians and GPU whisperers, come join us! We’re looking for supercomputing engineers to build the infrastructure behind real-time …

HardwareDGX agent

Cluster magicians and GPU whisperers, come join us! We’re looking for supercomputing engineers to build the infrastructure behind real-time interactive models, Tinker, and large-scale training: schedu

CME Group and Silicon Data announce a futures market for computing capacity, with contracts based on daily GPU benchmarks for on-demand rental rates (Tobias Burns/CNBC)

HardwareDGX agent

Tobias Burns / CNBC: CME Group and Silicon Data announce a futures market for computing capacity, with contracts based on daily GPU benchmarks for on-demand rental rates — A new futures market for sem

Did Jensen Huang catch conflict of interest disease from Sam?


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Hardware
DGX agent

Did Jensen Huang catch conflict of interest disease from Sam? HUANG FOUNDATION SIGNS GPU COMPUTE DEAL WITH COREWEAVE $NVDA proxy says the charitable foundation tied to Jensen and Lori Huang entered an

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

HardwareDGX agent

arXiv:2602.05391v2 Announce Type: replace Abstract: Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks.

Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis

HardwareDGX agent

arXiv:2511.08644v3 Announce Type: replace-cross Abstract: This paper presents a detailed comparative analysis of the performance of three major Python data manipulation libraries - Pandas, Polars, and

FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast

HardwareDGX agent

arXiv:2605.08314v1 Announce Type: cross Abstract: SVD-based Low-rank compression reduces transformer parameters and nominal FLOPs, but these savings often translate poorly into real LLM serving speedu

Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models

HardwareDGX agent

arXiv:2605.09681v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness,

Geometric 4D Stitching for Grounded 4D Generation

HardwareDGX agent

arXiv:2605.09984v1 Announce Type: cross Abstract: Recent 4D generation methods complete scene-level missing information using generative models and reconstruct the scene into radiance-based representa

Geometrically Approximated Modeling for Emitter-Centric Ray-Triangle Filtering in Arbitrarily Dynamic LiDAR Simulation

HardwareDGX agent

arXiv:2605.10457v1 Announce Type: cross Abstract: Real-time Light Detection And Ranging (LiDAR) simulation must find, per emitted ray, the closest intersecting triangle even in dynamic scenes containi

GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching

HardwareDGX agent

arXiv:2605.08243v1 Announce Type: cross Abstract: Synthesizing Mixed-Boolean Arithmetic (MBA) expressions from input-output examples is central to program deobfuscation and also useful for compiler op

How Imgix processes 8 billion images daily with G4 VMs powered by NVIDIA Blackwell

HardwareDGX agent

The modern web is extremely visual. People are busy and easily-distracted, and smart companies know they have just seconds to attract would-be customers with compelling images, videos, animations, and

Leveraging LLMs to Automate Energy-Aware Refactoring of Parallel Scientific Codes

HardwareDGX agent

arXiv:2505.02184v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for generating parallel scientific codes, with a primary focus on generating functionally correct

LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges

HardwareDGX agent

arXiv:2605.10807v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor i

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

HardwareDGX agent

arXiv:2605.10886v1 Announce Type: cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language

mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters

HardwareDGX agent

arXiv:2605.08300v1 Announce Type: cross Abstract: Manifold-Constrained Hyper-Connections (mHC) introduce a stability-motivated variant of multi stream residual mixing by constraining residual stream m

Model-Aware Tokenizer Transfer

HardwareDGX agent

arXiv:2510.21954v2 Announce Type: replace Abstract: Large Language Models (LLMs) are trained to support an increasing number of languages, yet their predefined tokenizers remain a bottleneck for adapt

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes

HardwareDGX agent

arXiv:2605.08913v1 Announce Type: cross Abstract: Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universall

Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning

HardwareDGX agent

arXiv:2605.09490v1 Announce Type: new Abstract: Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting

Novel GPU Boruta algorithms for feature selection from high-dimensional data

HardwareDGX agent

arXiv:2605.09950v1 Announce Type: cross Abstract: Most feature selection algorithms, especially wrapper methods, run inefficiently on CPU based platforms because of their high computational complexity

NVIDIA and SAP Bring Trust to Specialized Agents

HardwareDGX agent

Announced today at SAP Sapphire — where NVIDIA founder and CEO Jensen Huang joined SAP CEO Christian Klein’s keynote by video — SAP and NVIDIA’s expanded collaboration helps enterprises run specialize

Nvidia says that Jensen Huang is joining President Trump on his China trip; source: the president asked Huang to join after seeing media coverage of his absence (CNBC)

HardwareDGX agent

CNBC: Nvidia says that Jensen Huang is joining President Trump on his China trip; source: the president asked Huang to join after seeing media coverage of his absence — BEIJING — Nvidia CEO Jensen Hua

Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts

HardwareDGX agent

arXiv:2605.09055v1 Announce Type: cross Abstract: Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-

Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection

HardwareDGX agent

arXiv:2602.00318v2 Announce Type: replace-cross Abstract: The rise of bot accounts on social media poses significant risks to public discourse. To address this threat, modern bot detectors increasingl

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

HardwareDGX agent

arXiv:2605.09503v1 Announce Type: new Abstract: Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challengin

Qualcomm closed down 11.46% on Tuesday as chip stocks pull back from record AI-driven rally; Intel closed down 6.82%, Sandisk dropped 6%, and Micron 3.61% (Samantha Subin/CNBC)

HardwareDGX agent

Samantha Subin / CNBC: Qualcomm closed down 11.46% on Tuesday as chip stocks pull back from record AI-driven rally; Intel closed down 6.82%, Sandisk dropped 6%, and Micron 3.61% — Chip stocks dropped

SkillEvolver: Skill Learning as a Meta-Skill

HardwareDGX agent

arXiv:2605.10500v1 Announce Type: new Abstract: Agent skills today are static artifact: authored once -- by human curation or one-shot generation from parametric knowledge -- and then consumed unchang

Sources: Jensen Huang was left out of President Trump's China trip to avoid unwanted scrutiny and awkward conversations about the sale of Nvidia chips to China (Semafor)

HardwareDGX agent

Semafor: Sources: Jensen Huang was left out of President Trump's China trip to avoid unwanted scrutiny and awkward conversations about the sale of Nvidia chips to China — THE SCOOP — The Trump adminis

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

HardwareDGX agent

arXiv:2605.06356v2 Announce Type: replace Abstract: High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of t

The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls f…

HardwareDGX agent

The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much high

The EDA Primer: From RTL to Silicon

HardwareDGX agent

The EDA Primer covers the semiconductor design and manufacturing workflow, explaining how Electronic Design Automation tools transform Register Transfer Level (RTL) code into physical silicon through

Understand and Accelerate Memory Processing Pipeline for Disaggregated LLM Inference

HardwareDGX agent

arXiv:2603.29002v2 Announce Type: replace-cross Abstract: Modern large language models (LLMs) increasingly depends on efficient long-context processing and generation mechanisms, including sparse atte

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up ove…

HardwareDGX agent

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models,

11 May 2026

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference

HardwareDGX agent

arXiv:2605.07719v1 Announce Type: cross Abstract: Long-context inference increasingly operates over CPU-resident KV caches, either because decoding-time KV states exceed GPU memory capacity or because

CktFormalizer: Autoformalization of Natural Language into Circuit Representations

HardwareDGX agent

arXiv:2605.07782v1 Announce Type: new Abstract: LLMs can generate hardware descriptions from natural language specifications, but the resulting Verilog often contains width mismatches, combinational l

Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models

HardwareDGX agent

arXiv:2605.07194v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing method

DeepFedNAS: Efficient Hardware-Aware Architecture Adaptation for Heterogeneous IoT Federations via Pareto-Guided Supernet Training

HardwareDGX agent

arXiv:2601.15127v3 Announce Type: replace-cross Abstract: Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet

Direction-Preserving Number Representations

HardwareDGX agent

arXiv:2605.07662v1 Announce Type: new Abstract: Low-precision number formats are widely used in modern machine learning systems due to their efficiency. Accurate direction representation is key to the

Don't Learn the Shape: Forecasting Periodic Time Series by Rank-1 Decomposition

HardwareDGX agent

arXiv:2605.07222v1 Announce Type: new Abstract: How few parameters do we really need to forecast a periodic time series? An hourly electricity series, reshaped as a 24-row matrix with one column per d

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

HardwareDGX agent

arXiv:2605.07985v1 Announce Type: cross Abstract: Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, s

Everybody knows about the toilet maker Toto and MSG umami inventor Ajinomoto that are powering AI, but have you heard of the Taiwanese kitch…

HardwareDGX agent

This post likely discusses lesser-known Taiwanese companies in the kitchen or consumer appliance sector that play significant roles in powering AI infrastructure, drawing a parallel to how established

Frozen Backpropagation: Relaxing Weight Symmetry in Deep Spiking Neural Networks

HardwareDGX agent

arXiv:2505.13741v2 Announce Type: replace Abstract: Direct training of Spiking Neural Networks (SNNs) on neuromorphic hardware can greatly reduce energy costs compared to GPU-based training. However,

GATO: GPU-Accelerated and Batched Trajectory Optimization for Scalable Edge Model Predictive Control

HardwareDGX agent

arXiv:2510.07625v2 Announce Type: replace Abstract: While Model Predictive Control (MPC) delivers strong performance across robotics applications, solving the underlying (batches of) nonlinear traject

LLMSpace: Carbon Footprint Modeling for Large Language Model Inference on LEO Satellites

HardwareDGX agent

arXiv:2605.05615v2 Announce Type: replace Abstract: Large language models (LLMs) impose rapidly growing energy demands, creating an emerging energy and carbon crisis driven by large-scale inference. S

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

HardwareDGX agent

arXiv:2511.02805v2 Announce Type: replace-cross Abstract: LLM-based search agents often concatenate the full interaction history into the context, producing long and noisy inputs, and increasing compu

Nvidia embraces its role as an AI investor in 2026, committing 40B+ to equity investments including a 30B stake in OpenAI, 3.2B in Corning, and 2.1B in IREN (CNBC)

HardwareDGX agent

CNBC: Nvidia embraces its role as an AI investor in 2026, committing 40B+ to equity investments including a 30B stake in OpenAI, 3.2B in Corning, and 2.1B in IREN — Nvidia stepped on the gas last year

Physics-Based Flow Matching for Full-Field Prediction of Silicon Photonic Devices

HardwareDGX agent

arXiv:2605.06929v1 Announce Type: cross Abstract: Designing photonic integrated circuits requires accurate electromagnetic field simulations, which remain computationally expensive even for simple dev

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

HardwareDGX agent

arXiv:2605.06675v1 Announce Type: cross Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, mak

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement

HardwareDGX agent

arXiv:2605.06298v2 Announce Type: replace-cross Abstract: Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

HardwareDGX agent

arXiv:2605.07897v1 Announce Type: cross Abstract: Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded

SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers

HardwareDGX agent

arXiv:2603.02883v3 Announce Type: replace Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder ed

🧵 Slime: The Most Elegant & Comfortable RL Training Framework Ever A deep dive into why Slime redefines LLM RL training with clean architec…

HardwareDGX agent

🧵 Slime: The Most Elegant & Comfortable RL Training Framework Ever A deep dive into why Slime redefines LLM RL training with clean architecture & production-grade engineering ✨ Insights from Zhihu con

SOCKET: SOft Collision Kernel EsTimator for Sparse Attention

HardwareDGX agent

arXiv:2602.06283v2 Announce Type: replace Abstract: Exploiting sparsity during long-context inference is key to scaling large language models, as attention dominates the cost of autoregressive decodin

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache

HardwareDGX agent

arXiv:2605.06763v1 Announce Type: new Abstract: Sparse attention improves LLM inference efficiency by selecting a subset of key-value entries, but at the cost of potential accuracy degradation. In par

Sparser, Faster, Lighter Transformer Language Models

HardwareDGX agent

arXiv:2603.23198v2 Announce Type: replace-cross Abstract: Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, w

SpikingBrain: Spiking Brain-inspired Large Models

HardwareDGX agent

arXiv:2509.05276v4 Announce Type: replace-cross Abstract: Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence

Three things in AI to watch, according to a Nobel-winning economist

HardwareDGX agent

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. A few months before he was awarded the Nobel Prize in economic

Towards Billion-scale Multi-modal Biometric Search

HardwareDGX agent

arXiv:2605.07655v1 Announce Type: cross Abstract: Searching a multi-biometric database of a billion records for a country-level identity system requires pushing the limits of all aspects of a biometri

Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC

HardwareDGX agent

arXiv:2508.02001v2 Announce Type: replace-cross Abstract: Pervasive encryption makes large-scale labeling infeasible for traffic analysis, while security operations demand edge analysis to avert servi

XiYOLO: Energy-Aware Object Detection via Iterative Architecture Search and Scaling

HardwareDGX agent

arXiv:2605.06927v1 Announce Type: cross Abstract: Object detection on heterogeneous edge devices must satisfy strict energy, latency, and memory constraints while still providing reliable perception f

Zero-Shot Neural Network Evaluation with Sample-Wise Activation Patterns

HardwareDGX agent

arXiv:2605.07378v1 Announce Type: new Abstract: Zero-shot proxies, also known as training-free metrics, are widely adopted to reduce the computational overhead in neural network evaluation for scenari

← Previous
1…1920212223…29
Next →