AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,474 results
13 May 2026

ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference

HardwareDGX agent

arXiv:2605.11335v1 Announce Type: cross Abstract: Layerwise offloading reduces the GPU memory footprint of large diffusion transformer (DiT) inference by prefetching upcoming layers from host memory,

CME Group and Silicon Data to launch AI compute futures market

HardwareDGX agent

Silicon Data, the startup that provides market intelligence for artificial intelligence compute infrastructure, will provide the price indexes for a new futures market that will allow investors to hed

Efficient Remote KV Cache Reuse with GPU-native Video Codec

HardwareDGX agent

arXiv:2602.09725v3 Announce Type: replace-cross Abstract: Remote KV cache reuse fetches KV cache for identical contexts from remote storage, avoiding recomputation, accelerating LLM inference. While i

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Fast MoE Inference via Predictive Prefetching and Expert Replication

HardwareDGX agent

arXiv:2605.11537v1 Announce Type: new Abstract: The Mixture of Experts (MoE) architecture has become a fundamental building block in state-of-the-art large language models (LLMs), improving domain-spe

How the 'Facebook House' in Los Altos, Mark Zuckerberg's former residence, became a hub for Chinese AI talent, who are a key part of Silicon Valley's AI boom (Viola Zhou/Rest of World)

HardwareDGX agent

Viola Zhou / Rest of World: How the “Facebook House” in Los Altos, Mark Zuckerberg's former residence, became a hub for Chinese AI talent, who are a key part of Silicon Valley's AI boom — Chinese-born

NVIDIA, Ineffable Intelligence Team Up to Build the Future of Reinforcement Learning Infrastructure

HardwareDGX agent

Reinforcement-learning agents — AI systems that learn by trial and error — can convert computation into new knowledge. That’s the focus of a new engineering-level collaboration between NVIDIA and Inef

Recursive Superintelligence raises $650M to build self-improving AI models

HardwareDGX agent

Recursive Superintelligence Inc., a startup that hopes to develop self-improving artificial intelligence models, launched today with 650 million in funding. Alphabet Inc.’s GV fund and Greycroft led t

Red Hat and Intel spotlight scalable AI inference as enterprises move beyond the GPU gold rush

HardwareDGX agent

As companies move from testing AI to broader adoption, the biggest challenge is building scalable AI inference systems that perform without breaking the budget. The next wave of AI won’t be won on raw

Richard Socher's Recursive Superintelligence raised 650M+ from GV, Greycroft, Nvidia, AMD, and others at a 4B valuation to pursue 'recursive self-improvement' (Cade Metz/New York Times)

HardwareDGX agent

Cade Metz / New York Times: Richard Socher's Recursive Superintelligence raised 650M+ from GV, Greycroft, Nvidia, AMD, and others at a 4B valuation to pursue “recursive self-improvement” — Recursive S

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures

HardwareDGX agent

arXiv:2605.11999v1 Announce Type: cross Abstract: Power capping is the standard GPU energy lever in LLM serving, and it appears to work: throughput drops, power readings fall, and energy budgets are m

To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation

HardwareDGX agent

arXiv:2412.14461v4 Announce Type: replace Abstract: Unstructured text data annotation is foundational to management research. LLMs offer a cost-effective and scalable alternative to human annotation,

Transform Video Into Instantly Searchable, Actionable Intelligence with AI Agents and Skills

HardwareDGX agent

NVIDIA's video analytics AI agents analyze and process large volumes of video data through natural language tasks to provide critical insights , powered by vision language models, large language model

TriBand-BEV: Real-Time LiDAR-Only 3D Pedestrian Detection via Height-Aware BEV and High-Resolution Feature Fusion

HardwareDGX agent

arXiv:2605.12220v1 Announce Type: new Abstract: Safe autonomous agents and mobile robots need fast real time 3D perception, especially for vulnerable road users (VRUs) such as pedestrians. We introduc

12 May 2026

AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization

HardwareDGX agent

arXiv:2605.08692v1 Announce Type: cross Abstract: Post-training weight-only quantization to 4 bits is widely used to reduce the memory and compute costs of large language model inference. Existing PTQ

After studying 300 Leetcode Hards, solving every Jane Street puzzle from the Dwarkesh ads, and watching one Horace He lecture, he finally la…

HardwareDGX agent

After studying 300 Leetcode Hards, solving every Jane Street puzzle from the Dwarkesh ads, and watching one Horace He lecture, he finally landed the $400k annualized Jane Street internship. Unfortunat

CellDX AI Autopilot: Agent-Guided Training and Deployment of Pathology Classifiers

HardwareDGX agent

arXiv:2605.10362v1 Announce Type: new Abstract: Training AI models for computational pathology currently requires access to expensive whole-slide-image datasets, GPU infrastructure, deep expertise in

Cluster magicians and GPU whisperers, come join us! We’re looking for supercomputing engineers to build the infrastructure behind real-time …

HardwareDGX agent

Cluster magicians and GPU whisperers, come join us! We’re looking for supercomputing engineers to build the infrastructure behind real-time interactive models, Tinker, and large-scale training: schedu

CME Group and Silicon Data announce a futures market for computing capacity, with contracts based on daily GPU benchmarks for on-demand rental rates (Tobias Burns/CNBC)

HardwareDGX agent

Tobias Burns / CNBC: CME Group and Silicon Data announce a futures market for computing capacity, with contracts based on daily GPU benchmarks for on-demand rental rates — A new futures market for sem

Did Jensen Huang catch conflict of interest disease from Sam?

HardwareDGX agent

Did Jensen Huang catch conflict of interest disease from Sam? HUANG FOUNDATION SIGNS GPU COMPUTE DEAL WITH COREWEAVE $NVDA proxy says the charitable foundation tied to Jensen and Lori Huang entered an

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

HardwareDGX agent

arXiv:2602.05391v2 Announce Type: replace Abstract: Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks.

Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis

HardwareDGX agent

arXiv:2511.08644v3 Announce Type: replace-cross Abstract: This paper presents a detailed comparative analysis of the performance of three major Python data manipulation libraries - Pandas, Polars, and

FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast

HardwareDGX agent

arXiv:2605.08314v1 Announce Type: cross Abstract: SVD-based Low-rank compression reduces transformer parameters and nominal FLOPs, but these savings often translate poorly into real LLM serving speedu

Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models

HardwareDGX agent

arXiv:2605.09681v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness,

Geometric 4D Stitching for Grounded 4D Generation

HardwareDGX agent

arXiv:2605.09984v1 Announce Type: cross Abstract: Recent 4D generation methods complete scene-level missing information using generative models and reconstruct the scene into radiance-based representa

GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching

HardwareDGX agent

arXiv:2605.08243v1 Announce Type: cross Abstract: Synthesizing Mixed-Boolean Arithmetic (MBA) expressions from input-output examples is central to program deobfuscation and also useful for compiler op

Leveraging LLMs to Automate Energy-Aware Refactoring of Parallel Scientific Codes

HardwareDGX agent

arXiv:2505.02184v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for generating parallel scientific codes, with a primary focus on generating functionally correct

mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters

HardwareDGX agent

arXiv:2605.08300v1 Announce Type: cross Abstract: Manifold-Constrained Hyper-Connections (mHC) introduce a stability-motivated variant of multi stream residual mixing by constraining residual stream m

Model-Aware Tokenizer Transfer

HardwareDGX agent

arXiv:2510.21954v2 Announce Type: replace Abstract: Large Language Models (LLMs) are trained to support an increasing number of languages, yet their predefined tokenizers remain a bottleneck for adapt

Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning

HardwareDGX agent

arXiv:2605.09490v1 Announce Type: new Abstract: Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting

Novel GPU Boruta algorithms for feature selection from high-dimensional data

HardwareDGX agent

arXiv:2605.09950v1 Announce Type: cross Abstract: Most feature selection algorithms, especially wrapper methods, run inefficiently on CPU based platforms because of their high computational complexity

NVIDIA and SAP Bring Trust to Specialized Agents

HardwareDGX agent

Announced today at SAP Sapphire — where NVIDIA founder and CEO Jensen Huang joined SAP CEO Christian Klein’s keynote by video — SAP and NVIDIA’s expanded collaboration helps enterprises run specialize

Nvidia says that Jensen Huang is joining President Trump on his China trip; source: the president asked Huang to join after seeing media coverage of his absence (CNBC)

HardwareDGX agent

CNBC: Nvidia says that Jensen Huang is joining President Trump on his China trip; source: the president asked Huang to join after seeing media coverage of his absence — BEIJING — Nvidia CEO Jensen Hua

Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection

HardwareDGX agent

arXiv:2602.00318v2 Announce Type: replace-cross Abstract: The rise of bot accounts on social media poses significant risks to public discourse. To address this threat, modern bot detectors increasingl

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

HardwareDGX agent

arXiv:2605.09503v1 Announce Type: new Abstract: Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challengin

Qualcomm closed down 11.46% on Tuesday as chip stocks pull back from record AI-driven rally; Intel closed down 6.82%, Sandisk dropped 6%, and Micron 3.61% (Samantha Subin/CNBC)

HardwareDGX agent

Samantha Subin / CNBC: Qualcomm closed down 11.46% on Tuesday as chip stocks pull back from record AI-driven rally; Intel closed down 6.82%, Sandisk dropped 6%, and Micron 3.61% — Chip stocks dropped

SkillEvolver: Skill Learning as a Meta-Skill

HardwareDGX agent

arXiv:2605.10500v1 Announce Type: new Abstract: Agent skills today are static artifact: authored once -- by human curation or one-shot generation from parametric knowledge -- and then consumed unchang

Sources: Jensen Huang was left out of President Trump's China trip to avoid unwanted scrutiny and awkward conversations about the sale of Nvidia chips to China (Semafor)

HardwareDGX agent

Semafor: Sources: Jensen Huang was left out of President Trump's China trip to avoid unwanted scrutiny and awkward conversations about the sale of Nvidia chips to China — THE SCOOP — The Trump adminis

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

HardwareDGX agent

arXiv:2605.06356v2 Announce Type: replace Abstract: High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of t

The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls f…

HardwareDGX agent

The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much high

The EDA Primer: From RTL to Silicon

HardwareDGX agent

The EDA Primer covers the semiconductor design and manufacturing workflow, explaining how Electronic Design Automation tools transform Register Transfer Level (RTL) code into physical silicon through

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up ove…

HardwareDGX agent

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models,

11 May 2026

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference

HardwareDGX agent

arXiv:2605.07719v1 Announce Type: cross Abstract: Long-context inference increasingly operates over CPU-resident KV caches, either because decoding-time KV states exceed GPU memory capacity or because

Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models

HardwareDGX agent

arXiv:2605.07194v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing method

Code Generation and Conic Constraints for Model-Predictive Control on Microcontrollers with Conic-TinyMPC

ResearchDGX agent

arXiv:2403.18149v3 Announce Type: replace Abstract: Model-predictive control (MPC) is a state-of-the-art control method for constrained robotic systems, yet deployment on resource-limited hardware rem

Direction-Preserving Number Representations

HardwareDGX agent

arXiv:2605.07662v1 Announce Type: new Abstract: Low-precision number formats are widely used in modern machine learning systems due to their efficiency. Accurate direction representation is key to the

Don't Learn the Shape: Forecasting Periodic Time Series by Rank-1 Decomposition

HardwareDGX agent

arXiv:2605.07222v1 Announce Type: new Abstract: How few parameters do we really need to forecast a periodic time series? An hourly electricity series, reshaped as a 24-row matrix with one column per d

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

HardwareDGX agent

arXiv:2511.02805v2 Announce Type: replace-cross Abstract: LLM-based search agents often concatenate the full interaction history into the context, producing long and noisy inputs, and increasing compu

Nvidia embraces its role as an AI investor in 2026, committing 40B+ to equity investments including a 30B stake in OpenAI, 3.2B in Corning, and 2.1B in IREN (CNBC)

HardwareDGX agent

CNBC: Nvidia embraces its role as an AI investor in 2026, committing 40B+ to equity investments including a 30B stake in OpenAI, 3.2B in Corning, and 2.1B in IREN — Nvidia stepped on the gas last year

Physics-Based Flow Matching for Full-Field Prediction of Silicon Photonic Devices

HardwareDGX agent

arXiv:2605.06929v1 Announce Type: cross Abstract: Designing photonic integrated circuits requires accurate electromagnetic field simulations, which remain computationally expensive even for simple dev

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

HardwareDGX agent

arXiv:2605.06675v1 Announce Type: cross Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, mak

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement

HardwareDGX agent

arXiv:2605.06298v2 Announce Type: replace-cross Abstract: Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

HardwareDGX agent

arXiv:2605.07897v1 Announce Type: cross Abstract: Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded

🧵 Slime: The Most Elegant & Comfortable RL Training Framework Ever A deep dive into why Slime redefines LLM RL training with clean architec…

HardwareDGX agent

🧵 Slime: The Most Elegant & Comfortable RL Training Framework Ever A deep dive into why Slime redefines LLM RL training with clean architecture & production-grade engineering ✨ Insights from Zhihu con

SOCKET: SOft Collision Kernel EsTimator for Sparse Attention

HardwareDGX agent

arXiv:2602.06283v2 Announce Type: replace Abstract: Exploiting sparsity during long-context inference is key to scaling large language models, as attention dominates the cost of autoregressive decodin

Sparser, Faster, Lighter Transformer Language Models

HardwareDGX agent

arXiv:2603.23198v2 Announce Type: replace-cross Abstract: Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, w

Three things in AI to watch, according to a Nobel-winning economist

HardwareDGX agent

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. A few months before he was awarded the Nobel Prize in economic

Towards Billion-scale Multi-modal Biometric Search

HardwareDGX agent

arXiv:2605.07655v1 Announce Type: cross Abstract: Searching a multi-biometric database of a billion records for a country-level identity system requires pushing the limits of all aspects of a biometri

Zero-Shot Neural Network Evaluation with Sample-Wise Activation Patterns

HardwareDGX agent

arXiv:2605.07378v1 Announce Type: new Abstract: Zero-shot proxies, also known as training-free metrics, are widely adopted to reduce the computational overhead in neural network evaluation for scenari

10 May 2026

Arm, the UK and Apple

HardwareDGX agent

This article likely examines the relationship between Arm Holdings (the British semiconductor design company), the UK government, and Apple, possibly covering topics such as Apple's use of Arm-based c

I was talking to a room of senior accountants a couple months ago and 10% had OpenClaw installations. Of course there are far more non-users…

HardwareDGX agent

I was talking to a room of senior accountants a couple months ago and 10% had OpenClaw installations. Of course there are far more non-users and firms lag behind their people, but there is a sort of S

← Previous
1…2728293031…75
Next →