AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,446 results
20 May 2026

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

HardwareDGX agent

arXiv:2605.20179v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization an

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

HardwareDGX agent

arXiv:2605.19593v1 Announce Type: new Abstract: Modern deployments of Large Language Models (LLMs) increasingly require serving multiple models with diverse architectures, sizes, and specialization on

COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones

HardwareDGX agent

arXiv:2605.19138v1 Announce Type: cross Abstract: The scarcity of large-scale, high-quality demonstration data remains a bottleneck in scaling imitation learning for robotic manipulation. We present C

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
19 May 2026

LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models

HardwareDGX agent

arXiv:2605.17289v1 Announce Type: cross Abstract: Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execution to the

18 May 2026

TrainMover: An Interruption-Resilient Runtime for ML Training

HardwareDGX agent

arXiv:2412.12636v3 Announce Type: replace-cross Abstract: Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solution

15 May 2026

GenAI for Energy-Efficient and Interference-Aware Compressed Sensing of GNSS Signals on a Google Edge TPU

HardwareDGX agent

arXiv:2605.14839v1 Announce Type: new Abstract: Traditional methods for classifying global navigation satellite system (GNSS) jamming signals typically involve post-processing raw or spectral data str

11 May 2026

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

HardwareDGX agent

arXiv:2605.07985v1 Announce Type: cross Abstract: Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, s

GATO: GPU-Accelerated and Batched Trajectory Optimization for Scalable Edge Model Predictive Control

HardwareDGX agent

arXiv:2510.07625v2 Announce Type: replace Abstract: While Model Predictive Control (MPC) delivers strong performance across robotics applications, solving the underlying (batches of) nonlinear traject

LLMSpace: Carbon Footprint Modeling for Large Language Model Inference on LEO Satellites

HardwareDGX agent

arXiv:2605.05615v2 Announce Type: replace Abstract: Large language models (LLMs) impose rapidly growing energy demands, creating an emerging energy and carbon crisis driven by large-scale inference. S

Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC

HardwareDGX agent

arXiv:2508.02001v2 Announce Type: replace-cross Abstract: Pervasive encryption makes large-scale labeling infeasible for traffic analysis, while security operations demand edge analysis to avert servi

7 May 2026

Nvidia up 2.6% on news that xAI is flooding the market with 220,000 secondhand GPUs.

HardwareDGX agent

Nvidia's stock rose 2.6% following news that xAI is selling 220,000 used GPUs on the secondhand market, potentially indicating a shift in AI hardware demand or xAI's operational priorities. The report

When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI

HardwareDGX agent

arXiv:2605.03213v1 Announce Type: cross Abstract: Agentic AI systems, specifically LLM-driven agents that plan, invoke tools, maintain persistent memory, and delegate tasks to peer agents via protocol

4 May 2026

llamacpp on Apple Silicon, once configured correctly is really rock solid! You can throw anything at it and it will answer. Impressive!

HardwareDGX agent

Llamacpp, when properly configured on Apple Silicon hardware, demonstrates robust performance and reliability for running language models. The tool can handle varied input requests effectively, making

Most AI teams treat compute as a commodity. It's not.

HardwareDGX agent

AI compute resources should not be treated as interchangeable commodities, as different hardware configurations, providers, and architectures significantly impact training costs, inference latency, an

27 Apr 2026

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches

HardwareDGX agent

arXiv:2509.22166v4 Announce Type: replace-cross Abstract: The demand for efficient large language model (LLM) inference has intensified the focus on sparsification techniques. While semi-structured (N

22 Apr 2026

We're launching two specialized TPUs for the agentic era.

HardwareDGX agent

Google announced Ironwood, its seventh-generation TPU that is twice as power efficient as the previous generation, alongside specialized hardware designed to support the emerging agentic AI era. Ironw

20 Apr 2026

How Much Do GPU Clusters Really Cost?

HardwareDGX agent

This article from SemiAnalysis analyzes the total cost of ownership for GPU clusters, likely covering hardware expenses, infrastructure requirements, power consumption, and operational overhead beyond

NeuroMesh: A Unified Neural Inference Framework for Decentralized Multi-Robot Collaboration

HardwareDGX agent

arXiv:2604.15475v1 Announce Type: new Abstract: Deploying learned multi-robot models on heterogeneous robots remains challenging due to hardware heterogeneity, communication constraints, and the lack

17 Apr 2026

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers

HardwareDGX agent

arXiv:2509.23638v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memor

15 Apr 2026

Fast AI Model Partition for Split Learning over Edge Networks

HardwareDGX agent

arXiv:2507.01041v4 Announce Type: replace-cross Abstract: Split learning (SL) is a distributed learning paradigm that can enable computation-intensive artificial intelligence (AI) applications by part

14 Apr 2026

GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs

HardwareDGX agent

arXiv:2604.11659v1 Announce Type: cross Abstract: Fully homomorphic encryption (FHE) has recently attracted significant attention as both a cryptographic primitive and a systems challenge. Given the l

Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search

HardwareDGX agent

arXiv:2604.11109v1 Announce Type: cross Abstract: As high-performance computing and AI workloads become increasingly dependent on GPUs, maintaining high performance across rapidly evolving hardware ge

10 Apr 2026

LAsset: An LLM-assisted Security Asset Identification Framework for System-on-Chip (SoC) Verification

HardwareDGX agent

arXiv:2601.02624v2 Announce Type: replace-cross Abstract: The growing complexity of modern system-on-chip (SoC) and IP designs is making security assurance difficult day by day. One of the fundamental

QNAS: A Neural Architecture Search Framework for Accurate and Efficient Quantum Neural Networks

Model ReleasesDGX agent

arXiv:2604.07013v1 Announce Type: cross Abstract: Designing quantum neural networks (QNNs) that are both accurate and deployable on NISQ hardware is challenging. Handcrafted ansatze must balance expre

12 Aug 2026

CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

HardwareDGX agent

arXiv:2608.10506v1 Announce Type: cross Abstract: Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resourc

EvoMem: Memory-Augmented Evolution for Code Optimization

HardwareDGX agent

arXiv:2608.10795v1 Announce Type: new Abstract: Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may tran

11 Aug 2026

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation

HardwareDGX agent

arXiv:2608.08469v1 Announce Type: new Abstract: Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duple

DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Model ReleasesDGX agent

Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot o

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

Local AiDGX agent

arXiv:2608.08285v1 Announce Type: new Abstract: We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs

🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sa…

HardwareDGX agent

🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sachs, and KKR, to raise over $500 billion so its own customers

10 Aug 2026

Ising Acceleration for Multi-Robot Multi-Target Planning

ResearchDGX agent

arXiv:2608.06803v1 Announce Type: cross Abstract: Ising machines are emerging as promising hardware for combinatorial optimization. With recent advances in CMOS Ising technology, they are becoming att

Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Bat…

HardwareDGX agent

Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, High throughput prefill eng

6 Aug 2026

AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation

ResearchDGX agent

arXiv:2608.04502v1 Announce Type: cross Abstract: Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) language model

RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis

Model ReleasesDGX agent

arXiv:2602.11506v4 Announce Type: replace-cross Abstract: The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characteri

5 Aug 2026

SAMSEM -- A Generic and Scalable Approach for IC Metal Line Segmentation

ApplicationsDGX agent

arXiv:2603.16548v2 Announce Type: replace-cross Abstract: In light of globalized hardware supply chains, the assurance of hardware components has gained significant interest, particularly in cryptogra

31 Jul 2026

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

HardwareDGX agent

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a co

28 Jul 2026

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

HardwareDGX agent

arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next

24 Jul 2026

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

Model ReleasesDGX agent

arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchma

23 Jul 2026

PN-QNN: Harnessing Physical Noise as a Native Regularizer in Photonic Hybrid Quantum Neural Networks

Model ReleasesDGX agent

arXiv:2607.20045v1 Announce Type: cross Abstract: Physical noise in near-term quantum hardware is usually treated as a nuisance to suppress. We ask whether it can instead act as a hardware-native regu

ReRAM-aware Model Finetuning addressing I-V Non-linearity and Retention Errors

HardwareDGX agent

arXiv:2606.17471v3 Announce Type: replace Abstract: Traditional CPU, GPU, and NPU architectures are increasingly limited by the von Neumann bottleneck. While In-Memory Computing (IMC) using ReRAM cros

8 Jul 2026

BitLogic: Training Framework for Gradient-Based FPGA-Native Neural Networks

HardwareDGX agent

arXiv:2602.07400v2 Announce Type: replace Abstract: Gradient-based LUT- and logic-gate-based neural networks (LUTNet, LogicNets, DiffLogic, PolyLUT, NeuraLUT, WARP-LUT, DWN, LILogicNet, LightLUT) repl

DBNN: Neural Spike Classification Using a Deep Binarized Neural Network

HardwareDGX agent

arXiv:2607.05590v1 Announce Type: cross Abstract: Implantable brain-computer interfaces require on-node spike sorting to reduce telemetry bandwidth and power while maintaining reliable neural decoding

Life Cycle Assessment of Pre-training the Lucie 7B Open-Source Large Language Model on the Jean Zay Supercomputer

HardwareDGX agent

arXiv:2607.05408v1 Announce Type: cross Abstract: The environmental impact of training large language models (LLMs) is increasingly scrutinised, yet most published estimates focus on operational energ

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

HardwareDGX agent

arXiv:2607.06558v1 Announce Type: new Abstract: Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demons

7 Jul 2026

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems

HardwareDGX agent

arXiv:2607.02558v1 Announce Type: cross Abstract: As machine learning shifts from laboratory curiosity to critical infrastructure, the systems that sustain it span an extraordinary range, from sub-mil

3 Jul 2026

WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

HardwareDGX agent

arXiv:2607.02391v1 Announce Type: cross Abstract: Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requi

30 Jun 2026

CSAR: Containerized System Architecture for Robotics

HardwareDGX agent

arXiv:2606.30293v1 Announce Type: new Abstract: Robotic applications increasingly rely on distributed computational infrastructures that combine embedded devices, edge servers, and cloud resources. Th

26 Jun 2026

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

HardwareDGX agent

arXiv:2606.26453v1 Announce Type: new Abstract: We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating

10 Jun 2026

The Inference Alpha: Maximizing Frontier Models on AMD

IndustryDGX agent

This article discusses strategies for optimizing the performance of advanced AI frontier models when running on AMD hardware infrastructure. It likely covers deployment best practices, hardware config

Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA

AgentsDGX agent

arXiv:2606.11117v1 Announce Type: cross Abstract: Designing FPGA-based accelerators for modern artificial intelligence workloads requires exploring a large and complex hardware design space that invol

9 Jun 2026

LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load

Model ReleasesDGX agent

arXiv:2603.23640v2 Announce Type: replace-cross Abstract: Deploying large language models on-device for always-on personal agents demands sustained inference from hardware tightly constrained in power

Towards Automated Kernel Generation in the Era of LLMs

HardwareDGX agent

arXiv:2601.15727v3 Announce Type: replace Abstract: The performance of modern AI systems is fundamentally constrained by the quality of their underlying GPU kernels, which translate high-level algorit

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware

5 Jun 2026

What are the most capable LLM models I can run on my laptop?

Local AiDGX agent

A discussion on r/ollama exploring which high-performance LLM models can be effectively run locally on standard laptop hardware , likely covering model size comparisons, hardware requirements, and per

2 Jun 2026

Bit-Exact AI Inference Verification Without Performance Tradeoffs

HardwareDGX agent

arXiv:2606.00279v1 Announce Type: cross Abstract: Verifying claims about AI workloads is a pre- requisite for credible AI governance of covert adversaries (who comply with monitoring only when detecti

29 May 2026

Bastion: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting

HardwareDGX agent

arXiv:2605.29727v1 Announce Type: new Abstract: Block-diffusion drafters have recently emerged as a powerful alternative for speculative decoding by predicting multiple future-token distributions in a

26 May 2026

Inference Time Context Sparsity: Illusion or Opportunity?

HardwareDGX agent

arXiv:2605.24168v1 Announce Type: new Abstract: Sparsity has long been a central theme in LLM efficiency, but its role in context processing remains unresolved. As LLM workloads shift toward longer co

22 May 2026

ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration

HardwareDGX agent

arXiv:2605.22015v1 Announce Type: new Abstract: Diffusion Transformer (DiT) has emerged as a powerful model architecture for generating high-quality images and videos. In the case of video DiT, 3D Spa

12 May 2026

How Imgix processes 8 billion images daily with G4 VMs powered by NVIDIA Blackwell

HardwareDGX agent

The modern web is extremely visual. People are busy and easily-distracted, and smart companies know they have just seconds to attract would-be customers with compelling images, videos, animations, and

5 May 2026

A Semantic Autonomy Framework for VLM-Integrated Indoor Mobile Robots: Hybrid Deterministic Reasoning and Cross-Robot Adaptive Memory

HardwareDGX agent

arXiv:2605.02525v1 Announce Type: new Abstract: Autonomous indoor mobile robots can navigate reliably to metric coordinates using established frameworks such as ROS 2 Navigation 2, yet they lack the a

← Previous
1…45678…75
Next →