AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,450 results
Hardware

Efficient Training on Multiple Consumer GPUs with RoundPipe

DGX agent

arXiv:2604.27085v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on consumer-grade GPUs is highly cost-effective, yet constrained by limited GPU memory and slow PCIe intercon

hardwarearxiv-cs-ai
1 May 2026
Hardware
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

DGX agent

arXiv:2604.26968v1 Announce Type: cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving. Current

hardwarearxiv-cs-ai
1 May 2026
Hardware

AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices

DGX agent

arXiv:2604.25326v2 Announce Type: replace-cross Abstract: Speculative decoding enhances the inference efficiency of large language models (LLMs) by generating drafts using a small draft language model

hardwarearxiv-cs-ai
30 Apr 2026
Hardware

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving

DGX agent

arXiv:2604.26103v1 Announce Type: cross Abstract: All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneou

hardwarearxiv-cs-ai
30 Apr 2026
Hardware

Speed Up Unreal Engine NNE Inference with NVIDIA TensorRT for RTX Runtime

DGX agent

The TensorRT for RTX plugin provides a runtime for Unreal Engine's Neural Network Engine (NNE), enabling efficient deployment of AI models directly within real-time applications. Developers can see 1.

hardwarenvidia-developer
30 Apr 2026
Local Ai

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

DGX agent

arXiv:2603.15954v2 Announce Type: replace Abstract: Real-time AI experiences call for on-device large language models (OD-LLMs) optimized for efficient deployment on resource-constrained hardware. The

local-aiarxiv-cs-lg
29 Apr 2026
Hardware

PointTransformerX:Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms

DGX agent

arXiv:2604.24169v1 Announce Type: new Abstract: 3D point cloud perception remains tightly coupled to custom CUDA operators for spatial operations, limiting portability and efficiency on non-NVIDIA, AM

hardwarearxiv-cs-cv
28 Apr 2026
Hardware

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

DGX agent

arXiv:2505.02922v3 Announce Type: replace Abstract: Recent large language models (LLMs) are rapidly extending their context windows, yet inference throughput lags due to increasing GPU memory and band

hardwarearxiv-cs-lg
28 Apr 2026
Research

Bridging the Training-Deployment Gap: Gated Encoding and Multi-Scale Refinement for Efficient Quantization-Aware Image Enhancement

DGX agent

arXiv:2604.21743v1 Announce Type: new Abstract: Image enhancement models for mobile devices often struggle to balance high output quality with the fast processing speeds required by mobile hardware. W

researcharxiv-cs-ai
24 Apr 2026
Hardware

Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)

DGX agent

arXiv:2604.16395v2 Announce Type: replace-cross Abstract: Context retrieval systems for LLM inference face a critical challenge: high retrieval latency creates a fundamental tension between waiting fo

hardwarearxiv-cs-ai
24 Apr 2026
Hardware

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

DGX agent

arXiv:2412.03594v3 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly play an important role in a wide range of information processing and management tasks in industry. M

hardwarearxiv-cs-ai
23 Apr 2026
Model Releases

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because t…

DGX agent

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because the headlines are doing you a disservice: Elon Musk got on th

model-releasesgary-marcus--x
23 Apr 2026
Hardware

Stream-CQSA: Avoiding Out-of-Memory in Attention Computation via Flexible Workload Scheduling

DGX agent

arXiv:2604.20819v1 Announce Type: new Abstract: The scalability of long-context large language models is fundamentally limited by the quadratic memory cost of exact self-attention, which often leads t

hardwarearxiv-cs-lg
23 Apr 2026
Hardware

ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants

DGX agent

arXiv:2604.18616v1 Announce Type: cross Abstract: LLM-based coding agents can generate functionally correct GPU kernels, yet their performance remains far below hand-optimized libraries on critical co

hardwarearxiv-cs-ai
22 Apr 2026
Hardware

From Rainforests to Recycling Plants: 5 Ways NVIDIA AI Is Protecting the Planet

DGX agent

NVIDIA AI and accelerated computing are advancing sustainability, climate science and energy efficiency through five key applications. These include RecycleOS, an AI and robotics solution that helps r

hardwarenvidia-blog
22 Apr 2026
Hardware

Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning

DGX agent

arXiv:2412.00069v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) has garnered significant attention for its ability to scale up neural networks while utilizing the same or even fewer

hardwarearxiv-cs-cl
21 Apr 2026
Hardware

Enabling AI ASICs for Zero Knowledge Proof

DGX agent

arXiv:2604.17808v1 Announce Type: cross Abstract: Zero-knowledge proof (ZKP) provers remain costly because multi-scalar multiplication (MSM) and number-theoretic transforms (NTTs) dominate runtime as

hardwarearxiv-cs-cl
21 Apr 2026
Model Releases

MerLin: A Discovery Engine for Photonic and Hybrid Quantum Machine Learning

DGX agent

arXiv:2602.11092v2 Announce Type: replace Abstract: Identifying where quantum models may offer practical benefits in near term quantum machine learning (QML) requires moving beyond isolated algorithmi

model-releasesarxiv-cs-lg
21 Apr 2026
Hardware

RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts

DGX agent

arXiv:2510.04008v5 Announce Type: replace Abstract: Softmax Attention has a quadratic time complexity in sequence length, which becomes prohibitive to run at long contexts, even with highly optimized

hardwarearxiv-cs-lg
21 Apr 2026
Hardware

Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs

DGX agent

arXiv:2604.16715v1 Announce Type: cross Abstract: Graph foundation models have demonstrated remarkable adaptability across diverse downstream tasks through large-scale pretraining on graphs. However,

hardwarearxiv-cs-lg
21 Apr 2026
Research

Towards Real-Time ECG and EMG Modeling on mu NPUs

DGX agent

arXiv:2604.18067v1 Announce Type: new Abstract: The miniaturisation of neural processing units (NPUs) and other low-power accelerators has enabled their integration into microcontroller-scale wearable

researcharxiv-cs-lg
21 Apr 2026
Hardware

Web-Gewu: A Browser-Based Interactive Playground for Robot Reinforcement Learning

DGX agent

arXiv:2604.17050v1 Announce Type: new Abstract: With the rapid development of embodied intelligence, robotics education faces a dual challenge: high computational barriers and cumbersome environment c

hardwarearxiv-cs-ro
21 Apr 2026
Research

Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension

DGX agent

arXiv:2604.15769v1 Announce Type: cross Abstract: Spiking transformers achieve competitive accuracy with conventional transformers while offering 38-57imes energy efficiency on neuromorphic hardware,

researcharxiv-cs-ai
20 Apr 2026
Hardware

Scalable Posterior Uncertainty for Flexible Density-Based Clustering

DGX agent

arXiv:2603.03188v2 Announce Type: replace-cross Abstract: We introduce a novel framework for uncertainty quantification in clustering that combines martingale posterior distributions with density-base

hardwarearxiv-cs-lg
20 Apr 2026
Hardware

We’re launching Kimi K2.6 on Fireworks as a Day-0 launch partner! K2.5 was the base for standout models like @Cursor’s Composer 2 and was th…

DGX agent

We’re launching Kimi K2.6 on Fireworks as a Day-0 launch partner! K2.5 was the base for standout models like @Cursor’s Composer 2 and was the most popular model on our training platform. K2.6 on Firew

hardwarefireworks-ai--x
20 Apr 2026
Hardware

From Natural Language to PromQL: A Catalog-Driven Framework with Dynamic Temporal Resolution for Cloud-Native Observability

DGX agent

arXiv:2604.13048v1 Announce Type: cross Abstract: Modern cloud-native platforms expose thousands of time series metrics through systems like Prometheus, yet formulating correct queries in domain-speci

hardwarearxiv-cs-ai
17 Apr 2026
Hardware

Reference-Free Sampling-Based Model Predictive Control

DGX agent

arXiv:2511.19204v3 Announce Type: replace Abstract: We present a sampling-based model predictive control (MPC) framework that enables emergent locomotion without relying on handcrafted gait patterns o

hardwarearxiv-cs-ro
17 Apr 2026
Hardware

The Multi-GPU Era: Why Heterogeneous Compute Is Becoming Enterprise Standard

DGX agent

Heterogeneous computing, which combines multiple types of GPUs and processors, is becoming standard in enterprise environments to handle diverse workloads more efficiently than traditional homogeneous

hardwarevultr
16 Apr 2026
Local Ai

Cline with Ollama on a RTX4090 (24GRAM) and i9 with 64 GRAM

DGX agent

This Reddit post from r/ollama discusses a user's experience running Cline (an AI coding agent) with Ollama on a high-end local hardware setup consisting of an NVIDIA RTX 4090 with 24GB VRAM and an In

local-air-ollama
15 Apr 2026
Hardware

Roop Unleashed 4.3.1 not fully utilizing RTX 5070 Ti / 5080X (Low GPU/RAM usage)

DGX agent

This Reddit thread from r/StableDiffusion discusses a performance issue where Roop Unleashed version 4.3.1 — a face-swapping tool commonly used alongside Stable Diffusion — fails to fully utilize the

hardwarer-stablediffusion
15 Apr 2026
Local Ai

Running a 31B model locally made me realize how insane LLM infra actually is

DGX agent

A Reddit post from r/ollama in which a user shares their experience running a 31B parameter model locally using Ollama, reflecting on the surprisingly demanding hardware and infrastructure requirement

local-air-ollama
15 Apr 2026
Hardware

🆕 The Full Story of Notion AI https://latent.space/p/notion We're so excited to chat with @simonlast and @sarahmsachs about Notion's 'Token…

DGX agent

🆕 The Full Story of Notion AI https://latent.space/p/notion We're so excited to chat with @simonlast and @sarahmsachs about Notion's 'Token Town' - the crack team of AI Engineers and Model Behavior En

hardwareswyx--x
15 Apr 2026
Research

CUTEv2: Unified and Configurable Matrix Extension for Diverse CPU Architectures with Minimal Design Overhead

DGX agent

arXiv:2604.11615v1 Announce Type: cross Abstract: Matrix extensions have emerged as an essential feature in modern CPUs to address the surging demands of AI workloads. However, existing designs often

researcharxiv-cs-ai
14 Apr 2026
Safety

Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation

DGX agent

arXiv:2604.10170v1 Announce Type: cross Abstract: The growing complexity of visuomotor policies poses significant challenges for deployment with heterogeneous robotic hardware constraints. However, mo

safetyarxiv-cs-cv
14 Apr 2026
Hardware

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs

DGX agent

arXiv:2604.10539v1 Announce Type: cross Abstract: Key-Value (KV) cache plays a crucial role in accelerating inference in large language models (LLMs) by storing intermediate attention states and avoid

hardwarearxiv-cs-ai
14 Apr 2026
Hardware

NVIDIA Ising Introduces AI-Powered Workflows to Build Fault-Tolerant Quantum Systems

DGX agent

NVIDIA Ising is the world's first family of open-source quantum AI models, designed to help researchers and enterprises build quantum processors capable of running useful applications. The family span

hardwarenvidia-developer
14 Apr 2026
Hardware

SCNO: Spiking Compositional Neural Operator -- Towards a Neuromorphic Foundation Model for Nuclear PDE Solving

DGX agent

arXiv:2604.11625v1 Announce Type: cross Abstract: Neural operators have emerged as powerful surrogates for partial differential equation (PDE) solvers, yet they are typically trained as monolithic mod

hardwarearxiv-cs-ai
14 Apr 2026
Hardware

Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation

DGX agent

arXiv:2604.10180v1 Announce Type: cross Abstract: Disaggregation maps parts of an AI workload to different types of GPUs, offering a path to utilize modern heterogeneous GPU clusters. However, existin

hardwarearxiv-cs-lg
14 Apr 2026
Research

Vestibular reservoir computing

DGX agent

arXiv:2604.09943v1 Announce Type: new Abstract: Reservoir computing (RC) is a computational framework known for its training efficiency, making it ideal for physical hardware implementations. However,

researcharxiv-cs-lg
14 Apr 2026
Hardware

what's the best place to buy GPU server?

DGX agent

This Reddit thread from r/ollama discusses community recommendations for purchasing or renting GPU servers to run Ollama and local LLMs. It likely covers options ranging from dedicated GPU server prov

hardwarer-ollama
14 Apr 2026
Hardware

$200/month is enough to buy an H100 GPU for 6 hours every workday

DGX agent

Soumith Chintala shared a post highlighting that $200 per month is sufficient to rent access to an NVIDIA H100 GPU for approximately 6 hours every workday, making high-end AI compute more accessible t

hardwaresoumith-chintala--x
13 Apr 2026
Hardware

🫡 @aiDotEngineer in London was extremely good. Absolutely incredible speaker line up, great curation, escaped the Silicon Valley bubble mas…

DGX agent

"**AI Engineer Europe (London) — Conference Recap**

hardwareswyx--x
10 Apr 2026
Hardware

Behind the Analysis with Google Cloud and Team USA: Architecting AI infrastructure for U.S. Winter Olympians

DGX agent

In freeskiing and snowboarding, traditional video replay shows you what happened during a complex aerial maneuver, but it fails to explain the physics of how it was possible. At the speed of the sport

hardwaregoogle-cloud-ai
10 Apr 2026
Hardware

That’s a wrap on HumanX. Custom comics, hats, a happy hour with @getmetronome & @nvidia, and two sessions on what actually matters for AI-na…

DGX agent

Together Computer (Together AI) participated in HumanX 2026, a major AI conference held April 6–9 in San Francisco, where they hosted activations including custom comics, branded hats, and a happy ...

hardwaretogether-ai--x
10 Apr 2026
Hardware

Cut Checkpoint Costs with About 30 Lines of Python and NVIDIA nvCOMP

DGX agent

Training LLMs requires periodic checkpoints — full snapshots of model weights, optimizer states, and gradients — whose storage costs can reach $200,000/month for a 405B model on 128 NVIDIA DGX B200...

hardwarenvidia-developer
9 Apr 2026
Hardware

How to Accelerate Protein Structure Prediction at Proteome-Scale

DGX agent

NVIDIA, Google DeepMind, EMBL-EBI, and Seoul National University collaborated to extend the AlphaFold Protein Structure Database (AFDB) beyond monomeric structures to proteome-scale quaternary stru...

hardwarenvidia-developer
9 Apr 2026
Hardware

Running Large-Scale GPU Workloads on Kubernetes with Slurm

DGX agent

NVIDIA's open-source project **Slinky** (developed by SchedMD, now part of NVIDIA) enables organizations to run full Slurm clusters directly on Kubernetes infrastructure by managing the complete li...

hardwarenvidia-developer
9 Apr 2026
Hardware

Integrate Physical AI Capabilities into Existing Apps with NVIDIA Omniverse Libraries

DGX agent

NVIDIA has introduced a modular, library-based architecture for Omniverse, exposing core components—RTX rendering, PhysX-based simulation, and data storage pipelines—as standalone, headless-first C...

hardwarenvidia-developer
8 Apr 2026
← Previous
1…1011121314…93
Next →