AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,450 results
5 May 2026

A Semantic Autonomy Framework for VLM-Integrated Indoor Mobile Robots: Hybrid Deterministic Reasoning and Cross-Robot Adaptive Memory

HardwareDGX agent

arXiv:2605.02525v1 Announce Type: new Abstract: Autonomous indoor mobile robots can navigate reliably to metric coordinates using established frameworks such as ROS 2 Navigation 2, yet they lack the a

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning

HardwareDGX agent

arXiv:2602.13595v2 Announce Type: replace Abstract: Neural scaling laws provide a predictable recipe for AI advancement: reducing numerical precision should linearly improve computational efficiency a

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

HardwareDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.02262v1 Announce Type: cross Abstract: Recently, video language models (VLMs) have been applied in various fields. However, the visual token sequence of the VLM is too long, which may cause

1 May 2026

EdgeFM: Efficient Edge Inference for Vision-Language Models

HardwareDGX agent

arXiv:2604.27476v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated strong applicability in edge industrial applications, yet their deployment remains severely constrained

EdgeSpike: Spiking Neural Networks for Low-Power Autonomous Sensing in Edge IoT Architectures

Model ReleasesDGX agent

arXiv:2604.27004v1 Announce Type: cross Abstract: We propose EdgeSpike, a co-designed spiking neural network (SNN) framework for autonomous low-power sensing in edge Internet of Things (IoT) architect

29 Apr 2026

asRoBallet: Closing the Sim2Real Gap via Friction-Aware Reinforcement Learning for Underactuated Spherical Dynamics

Model ReleasesDGX agent

arXiv:2604.24916v1 Announce Type: new Abstract: We introduce asRoBallet, to the best of our knowledge, the first successful deployment of reinforcement learning (RL) on a humanoid ballbot hardware. Hi

@ClementDelangue I got mine too yesterday, fitting nicely next to my DGX Spark.

Local AiDGX agent

Clement Delangue, CEO of Hugging Face, posted about receiving a piece of hardware that fits alongside a DGX Spark system, likely referring to an AI accelerator or GPU device. The post suggests a perso

28 Apr 2026

Latent Inter-Frame Pruning: A Training-Free Method Bridging Traditional Video Compression and Modern Diffusion Transformers for Efficient Generation

HardwareDGX agent

arXiv:2604.23858v1 Announce Type: new Abstract: Video generation, while capable of generating realistic videos, is computationally expensive and slow, prohibiting real-time applications. In this paper

27 Apr 2026

HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference

ApplicationsDGX agent

arXiv:2604.22293v1 Announce Type: cross Abstract: Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations dir

25 Apr 2026

DAVIS, APRIL 25, 2026 — InferenceX has added DeepSeekv4 for @vllm_project 's day 0 support for GB200 disagg! Great work to @flowpow123 @roge…

HardwareDGX agent

DAVIS, APRIL 25, 2026 — InferenceX has added DeepSeekv4 for @vllm_project 's day 0 support for GB200 disagg! Great work to @flowpow123 @rogerw0108 @NVIDIAAIDev @inferact for the fast support and engin

23 Apr 2026

A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing

HardwareDGX agent

arXiv:2604.20129v1 Announce Type: new Abstract: The Synergistic Collapse occurs when scaling beyond 100 agents causes superlinear performance degradation that individual optimizations cannot prevent.

Amortized Vine Copulas for High-Dimensional Density and Information Estimation

HardwareDGX agent

arXiv:2604.20568v1 Announce Type: new Abstract: Modeling high-dimensional dependencies while keeping likelihoods tractable remains challenging. Classical vine-copula pipelines are interpretable but ca

22 Apr 2026

'Funny and distressingly realistic...propelled by awesome characters and inventive twists”— @andyweirauthor Silicon Valley invents the time …

HardwareDGX agent

'Funny and distressingly realistic...propelled by awesome characters and inventive twists”— @andyweirauthor Silicon Valley invents the time machine in my upcoming book PARADOX INC, now available for p

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

HardwareDGX agent

arXiv:2604.19129v1 Announce Type: new Abstract: Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment model

SpikeMLLM: Spike-based Multimodal Large Language Models via Modality-Specific Temporal Scales and Temporal Compression

HardwareDGX agent

arXiv:2604.18610v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but incur substantial computational overhead and energy consumption during

21 Apr 2026

A look at Apple's management changes and several key executives' futures; sources say Mike Rockwell has considered leaving or taking an advisory role next year (Mark Gurman/Bloomberg)

IndustryDGX agent

Mark Gurman / Bloomberg: A look at Apple's management changes and several key executives' futures; sources say Mike Rockwell has considered leaving or taking an advisory role next year — John Ternus,

AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation

HardwareDGX agent

arXiv:2604.18348v1 Announce Type: new Abstract: Video diffusion transformers (DiTs) suffer from prohibitive inference latency due to quadratic attention complexity. Existing sparse attention methods e

Framework’s first eGPUs turn its laptop into a desktop PC

HardwareDGX agent

Remember when Framework made the first laptop where you can easily upgrade its entire internal video card in three minutes flat? The company's getting into the external graphics game, too. As promised

Real-Time Structural Detection for Indoor Navigation from 3D LiDAR Using Bird's-Eye-View Images

HardwareDGX agent

arXiv:2603.19830v2 Announce Type: replace Abstract: Efficient structural perception is essential for mapping and autonomous navigation on resource-constrained robots. Existing 3D methods are computati

20 Apr 2026

CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment

HardwareDGX agent

arXiv:2604.15665v1 Announce Type: new Abstract: Markerless 3D movement analysis from monocular video enables accessible biomechanical assessment in clinical and sports settings. However, most research

Silicon Valley has forgotten what normal people want

HardwareDGX agent

One of the most mortifying things about knowing a lot of techies is listening to them tell me excitedly about some very important discovery that they believe they have made. Recently, I ran into an ac

Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge

HardwareDGX agent

arXiv:2604.15357v1 Announce Type: cross Abstract: Precise estimation of model inference latency is crucial for time-critical mobile edge applications, enabling devices to calculate latency margins aga

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective

HardwareDGX agent

arXiv:2511.00739v3 Announce Type: replace Abstract: Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on

17 Apr 2026

At SemiAnalysis, we're quite tired of the minimalist style webapps and landing pages that have become so commonplace lately. Today, we're in…

HardwareDGX agent

At SemiAnalysis, we're quite tired of the minimalist style webapps and landing pages that have become so commonplace lately. Today, we're introducing Minecraft mode on inferencex dot com so that you c

16 Apr 2026

Automatic Charge State Tuning of 300 mm FDSOI Quantum Dots Using Neural Network Segmentation of Charge Stability Diagram

HardwareDGX agent

arXiv:2604.13662v1 Announce Type: cross Abstract: Tuning of gate-defined semiconductor quantum dots (QDs) is a major bottleneck for scaling spin qubit technologies. We present a deep learning (DL) dri

15 Apr 2026

AMD Instinct™ GPU Preemptible Instances at Vultr

HardwareDGX agent

Vultr offers preemptible cloud instances powered by AMD Instinct GPUs, providing a cost-effective option for running GPU-accelerated workloads such as AI/ML training, inference, and high-performance c

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation

HardwareDGX agent

arXiv:2604.12798v1 Announce Type: cross Abstract: FlashAttention-style online softmax enables exact attention computation with linear memory by streaming score tiles through on-chip memory and maintai

14 Apr 2026

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows

HardwareDGX agent

arXiv:2604.09611v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in applications forming multi-request workflows like document summarization, search-based copilots,

From Decision Trees to Boolean Logic: A Fast and Unified SHAP Algorithm

HardwareDGX agent

arXiv:2511.09376v2 Announce Type: replace Abstract: SHapley Additive exPlanations (SHAP) is a key tool for interpreting decision tree ensembles by assigning contribution values to features. It is wide

Lifetime-Aware Design for Item-Level Intelligence at the Extreme Edge

HardwareDGX agent

arXiv:2509.08193v2 Announce Type: replace-cross Abstract: We present FlexiFlow, a lifetime-aware design framework for item-level intelligence (ILI) where computation is integrated directly into dispos

Simulator Adaptation for Sim-to-Real Learning of Legged Locomotion via Proprioceptive Distribution Matching

Model ReleasesDGX agent

arXiv:2604.11090v1 Announce Type: new Abstract: Simulation trained legged locomotion policies often exhibit performance loss on hardware due to dynamics discrepancies between the simulator and the rea

11 Aug 2026

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

HardwareDGX agent

arXiv:2608.08256v1 Announce Type: new Abstract: Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but

Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence

HardwareDGX agent

arXiv:2608.08761v1 Announce Type: cross Abstract: In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturin

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

HardwareDGX agent

arXiv:2607.18171v2 Announce Type: replace Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose effici

What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload

HardwareDGX agent

arXiv:2608.08287v1 Announce Type: new Abstract: GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement

10 Aug 2026

Beyond Visibility: Real-Time Surface Accessibility Fields from Sparse LiDAR

HardwareDGX agent

arXiv:2608.06412v1 Announce Type: cross Abstract: Understanding which surfaces in a scene are physically accessible to a given tool is fundamental for robotic interaction, yet 3D perception systems ty

Introducing the Developer Device Platform for agentic mobile app development

Model ReleasesDGX agent

Most enterprises connect with their customers through a device. Whether it’s using a mobile app to order a product, contact customer service, view content, or manage their account, the customer experi

6 Aug 2026

SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System

HardwareDGX agent

arXiv:2608.05033v1 Announce Type: cross Abstract: Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the in

5 Aug 2026

A Deployment-Friendly Foundational Framework for Efficient Computational Pathology

HardwareDGX agent

arXiv:2602.14010v2 Announce Type: replace-cross Abstract: Pathology foundation models (PFMs) generalize well across computational pathology tasks but remain costly for gigapixel whole-slide image anal

4 Aug 2026

SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits

HardwareDGX agent

arXiv:2608.00859v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. This in

Thermalizing Stochastic Programs

ResearchDGX agent

arXiv:2608.01615v1 Announce Type: cross Abstract: We present a set of tools for mapping general stochastic programs to thermodynamic hardware designed for energy-efficient stochastic sampling. Given a

3 Aug 2026

'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

Model ReleasesDGX agent

I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th

DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config

Model ReleasesDGX agent

Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results for this engine. Why bothe

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

HardwareDGX agent

arXiv:2607.28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below pe

How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure

HardwareDGX agent

A combined architecture using KAI Scheduler and vCluster enables multiple teams to run fully isolated Kubernetes tenant clusters on a shared GPU node, each with its own control plane, RBAC, CRDs, and

Topology-Aware Data Movement for Disaggregated GPU Inference

HardwareDGX agent

arXiv:2607.28633v1 Announce Type: cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate

1 Aug 2026

Vivid depiction of a dark scenario for Nvidia. “Indeed, CEO Jensen Huang has been so aggressive in funding the AI ecosystem that the company…

HardwareDGX agent

Vivid depiction of a dark scenario for Nvidia. “Indeed, CEO Jensen Huang has been so aggressive in funding the AI ecosystem that the company is functioning almost like a bank; with annual free cash fl

31 Jul 2026

Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

HardwareDGX agent

The article shows that dense‑attention performance in long‑context inference is governed by group size (query heads per KV head), head dimension, and sequence length, with prefill being compute‑bound

FasTac: A Curved Multispectral Vision-Based Tactile Sensor for High-Speed High-Precision 3D Shape and Force Perception

HardwareDGX agent

arXiv:2607.28416v1 Announce Type: new Abstract: Curved tactile fingertips for dexterous manipulation must resolve fine contact geometry, distinguish normal and tangential loads, and capture transient

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

Model ReleasesDGX agent

arXiv:2607.27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specia

QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction

Model ReleasesDGX agent

arXiv:2607.28422v1 Announce Type: new Abstract: Fault-tolerant quantum computing (FTQC) relies on quantum error correction to suppress physical errors and preserve logical information at scale. In pra

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

ApplicationsDGX agent

arXiv:2607.28418v1 Announce Type: cross Abstract: Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver

30 Jul 2026

Run High-Performance Core Math at Scale with NVIDIA nvmath-python

HardwareDGX agent

NVIDIA **nvmath‑python v1.0** is a Python library that wraps CUDA‑X and NVPL math libraries (cuFFT, cuBLASLt, cuDSS, cuSPARSE, cuTENSOR, cuBLASMp) to give NumPy, CuPy, and PyTorch users GPU‑accelerate

29 Jul 2026

When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning

SafetyDGX agent

arXiv:2607.25593v1 Announce Type: new Abstract: Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and un

28 Jul 2026

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

HardwareDGX agent

arXiv:2607.24148v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deploy

Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

HardwareDGX agent

arXiv:2607.24741v1 Announce Type: cross Abstract: Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distri

Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

HardwareDGX agent

arXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr

Multi-primitive in-memory computing for Monte Carlo tree search

HardwareDGX agent

arXiv:2607.22869v1 Announce Type: cross Abstract: Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge d

Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems

HardwareDGX agent

arXiv:2607.22714v1 Announce Type: cross Abstract: Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive plat

Rendering on Real Silicon: GPU Render-Timing as a Passive, AI-Resistant CAPTCHA Signal

HardwareDGX agent

arXiv:2607.23389v1 Announce Type: cross Abstract: Conventional CAPTCHAs pose puzzles that modern AI systems increasingly solve, while behavioral and cryptographic-attestation defenses carry privacy or

← Previous
1…56789…75
Next →