AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,450 results
29 May 2026

How Together AI built the world’s fastest speech-to-text stack

HardwareDGX agent

Together AI developed an optimized speech-to-text system focused on achieving the fastest processing speeds through technical innovations in their inference stack and model optimization. The approach

Q-ANCHOR: Federated Quantum Learning with ZNE-guided Correction

Local AiDGX agent

arXiv:2605.30075v1 Announce Type: new Abstract: Quantum Federated Learning (QFL) offers a promising framework to train quantum models across distributed clients while keeping data strictly local. Due

SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations

AgentsDGX agent

arXiv:2605.30345v1 Announce Type: new Abstract: Printed circuit board (PCB) schematic design defines nearly all electronic hardware, but it remains manual and expertise-intensive. While generative AI

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

What is the best used or refurbished laptop with GPU for open source Imege generation?

HardwareDGX agent

This Reddit discussion in the StableDiffusion community addresses recommendations for affordable, used or refurbished laptops equipped with GPUs suitable for running open-source image generation model

28 May 2026

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

Model ReleasesDGX agent

arXiv:2605.27472v1 Announce Type: cross Abstract: Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assert

Official @NVIDIAAI GLM5.1-NVFP4 spotted on @huggingface 🤩 https://huggingface.co/nvidia/GLM-5.1-NVFP4

HardwareDGX agent

NVIDIA has released GLM-5.1-NVFP4, a quantized version of the GLM-5.1 model, now available on Hugging Face. The model appears to use NVFP4 (NVIDIA's floating-point 4-bit) quantization format, designed

26 May 2026

Parallel Differentiable Reachability for Learning and Planning with Certified Neural Dynamics and Controllers

HardwareDGX agent

arXiv:2605.25346v1 Announce Type: cross Abstract: Neural network (NN) dynamics models and control policies achieve strong performance in robotics, but providing sound guarantees under uncertainty rema

SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling

HardwareDGX agent

arXiv:2605.24016v1 Announce Type: cross Abstract: Diffusion inference remains costly for edge deployment, yet existing accelerators focus almost exclusively on score networks because standard drift is

25 May 2026

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

HardwareDGX agent

arXiv:2605.22840v1 Announce Type: cross Abstract: How much thinking can a civilisation do? Kardashev's (1964) typology ranks civilisations by total power: planetary (Type I, ~10^16 W), stellar (Type I

20 May 2026

Bilateral Teleoperation with Compliant 6-DOF Pose-and-Force Sensing

ResearchDGX agent

arXiv:2605.19255v1 Announce Type: new Abstract: Existing bilateral teleoperation platforms still rely on costly rigid six-axis force/torque sensors, tightly coupled leader-follower hardware, and kiloh

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

HardwareDGX agent

arXiv:2605.19269v1 Announce Type: new Abstract: Transformer training systems are built around dense linear algebra, yet a nontrivial fraction of end-to-end time is spent on surrounding memory-bound op

Receptogenesis in a Vascularized Robotic Embodiment

ResearchDGX agent

arXiv:2603.09473v2 Announce Type: replace Abstract: Equipping robotic systems with the capacity to generate extit{ex novo} hardware during operation extends control of physical adaptability. Unlike mo

Soft Learning

HardwareDGX agent

arXiv:2605.18889v1 Announce Type: cross Abstract: Modern machine learning forces practitioners to choose between powerful but expensive deep networks and fast but limited classical algorithms. Here we

19 May 2026

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference

HardwareDGX agent

arXiv:2605.17164v1 Announce Type: cross Abstract: Deploying large-scale LLM training and inference with optimal performance is exceptionally challenging due to a complex design space of parallelism st

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

HardwareDGX agent

arXiv:2602.06694v2 Announce Type: replace Abstract: Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to effi

Stable Audio 3

HardwareDGX agent

arXiv:2605.17991v1 Announce Type: cross Abstract: Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models c

18 May 2026

A3D: Agentic AI flow for autonomous Accelerator Design

Model ReleasesDGX agent

arXiv:2605.15237v1 Announce Type: cross Abstract: Accelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite adva

15 May 2026

An Amortized Efficiency Threshold for Comparing Neural and Heuristic Solvers in Combinatorial Optimization

HardwareDGX agent

arXiv:2605.14624v1 Announce Type: cross Abstract: A common critique of neural combinatorial-optimization solvers is that they are less energy-efficient than CPU metaheuristics, given the operational e

Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI

HardwareDGX agent

arXiv:2605.14665v1 Announce Type: new Abstract: Legal reasoning is not semantic similarity search. A court judgment encodes constrained symbolic reasoning: precedent propagation, procedural state tran

FALO: Fast and Accurate LiDAR 3D Object Detection on Resource-Constrained Devices

HardwareDGX agent

arXiv:2506.04499v2 Announce Type: replace Abstract: Existing LiDAR 3D object detection methods predominantely rely on sparse convolutions and/or transformers, which can be challenging to run on resour

Towards Robotic Dexterous Hand Intelligence: A Survey

ResearchDGX agent

arXiv:2605.13925v1 Announce Type: new Abstract: Robotic dexterous hands are central to contact-rich manipulation, with rapid progress driven by advances in hardware, sensing, control, simulation, and

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference

HardwareDGX agent

arXiv:2605.14844v1 Announce Type: cross Abstract: We introduce XFP, a dynamic weight quantizer for LLM inference that inverts the conventional workflow: the operator specifies reconstruction quality f

14 May 2026

Geometric Autoencoder Priors for Bayesian Inversion: Learn First Observe Later

HardwareDGX agent

arXiv:2509.19929v4 Announce Type: replace-cross Abstract: Uncertainty Quantification (UQ) is paramount for inference in engineering. A common inference task is to recover full-field information of phy

13 May 2026

Cerebras — Faster Tokens Please

HardwareDGX agent

Cerebras, a company specializing in AI accelerators and wafer-scale computing systems, is discussed in terms of its approaches to improving token generation speed in large language models, which is cr

MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces

HardwareDGX agent

arXiv:2605.11333v1 Announce Type: cross Abstract: The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed ma

12 May 2026

Geometrically Approximated Modeling for Emitter-Centric Ray-Triangle Filtering in Arbitrarily Dynamic LiDAR Simulation

HardwareDGX agent

arXiv:2605.10457v1 Announce Type: cross Abstract: Real-time Light Detection And Ranging (LiDAR) simulation must find, per emitted ray, the closest intersecting triangle even in dynamic scenes containi

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

HardwareDGX agent

arXiv:2605.10886v1 Announce Type: cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes

HardwareDGX agent

arXiv:2605.08913v1 Announce Type: cross Abstract: Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universall

Understand and Accelerate Memory Processing Pipeline for Disaggregated LLM Inference

HardwareDGX agent

arXiv:2603.29002v2 Announce Type: replace-cross Abstract: Modern large language models (LLMs) increasingly depends on efficient long-context processing and generation mechanisms, including sparse atte

11 May 2026

Everybody knows about the toilet maker Toto and MSG umami inventor Ajinomoto that are powering AI, but have you heard of the Taiwanese kitch…

HardwareDGX agent

This post likely discusses lesser-known Taiwanese companies in the kitchen or consumer appliance sector that play significant roles in powering AI infrastructure, drawing a parallel to how established

SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers

HardwareDGX agent

arXiv:2603.02883v3 Announce Type: replace Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder ed

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache

HardwareDGX agent

arXiv:2605.06763v1 Announce Type: new Abstract: Sparse attention improves LLM inference efficiency by selecting a subset of key-value entries, but at the cost of potential accuracy degradation. In par

SpikingBrain: Spiking Brain-inspired Large Models

HardwareDGX agent

arXiv:2509.05276v4 Announce Type: replace-cross Abstract: Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence

XiYOLO: Energy-Aware Object Detection via Iterative Architecture Search and Scaling

HardwareDGX agent

arXiv:2605.06927v1 Announce Type: cross Abstract: Object detection on heterogeneous edge devices must satisfy strict energy, latency, and memory constraints while still providing reliable perception f

8 May 2026

Great collab with @SakanaAILabs on an #ICML26 paper about sparse transformer kernels + formats optimized for modern NVIDIA GPU execution. • …

HardwareDGX agent

Great collab with @SakanaAILabs on an #ICML26 paper about sparse transformer kernels + formats optimized for modern NVIDIA GPU execution. • TwELL sparse packing • Fused CUDA kernels • 20%+ inference/t

With faster node startup for GKE, say goodbye to cold-start latency

HardwareDGX agent

We’ve rolled out a significant update to Google Kubernetes Engine (GKE) that solves one of the most annoying problems in cloud infrastructure: cold start latency. GKE now has up to 4x faster node star

7 May 2026

Achieving Peak System and Workload Efficiency on NVIDIA GB200 NVL72 with Slurm Block Scheduling

HardwareDGX agent

This article discusses optimization techniques for maximizing efficiency on NVIDIA's GB200 NVL72 system using Slurm block scheduling, a job scheduling approach designed to improve resource utilization

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs

HardwareDGX agent

arXiv:2605.04357v1 Announce Type: cross Abstract: The usage of large language models (LLMs) has grown increasingly fragmented, with no single model dominating. Meanwhile, cloud providers offer a wide

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism

HardwareDGX agent

arXiv:2605.05049v1 Announce Type: cross Abstract: Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE mo

6 May 2026

Nvidia partners with Corning to boost the supply of optical network components

HardwareDGX agent

Nvidia Corp. will help publicly traded glass maker Corning Inc. boost the rate at which it produces parts for optical data center networks. The partnership, which the companies announced today, will s

VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU

HardwareDGX agent

arXiv:2605.01352v1 Announce Type: cross Abstract: GPU-based simulation environments for embodied AI interleave physics simulation (CUDA) and photorealistic rendering (Vulkan) on a single device. We ob

5 May 2026

A Synthesizable RTL Implementation of Predictive Coding Networks

Local AiDGX agent

arXiv:2603.18066v2 Announce Type: replace-cross Abstract: Backpropagation has enabled modern deep learning but is difficult to realize as an online, fully distributed hardware learning system due to g

aerial-autonomy-stack -- a Faster-than-real-time, Autopilot-agnostic, ROS2 Framework to Simulate and Deploy Perception-based Drones

HardwareDGX agent

arXiv:2602.07264v2 Announce Type: replace Abstract: Unmanned aerial vehicles are rapidly transforming multiple applications, from agricultural and infrastructure monitoring to logistics and defense. I

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization

HardwareDGX agent

arXiv:2602.02958v4 Announce Type: replace Abstract: Despite rapid progress in autoregressive video diffusion, an emerging system algorithm bottleneck limits both deployability and generation capabilit

Silicon Valley bets $200M on AI data centers floating in the ocean

HardwareDGX agent

Panthalassa, a US startup, is developing floating data centers powered by ocean waves and cooled by seawater to address AI infrastructure's growing energy demands. Peter Thiel led a $140 million inves

4 May 2026

DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference

HardwareDGX agent

arXiv:2605.00174v1 Announce Type: cross Abstract: Video and image streaming on edge devices requires low latency. To address this, Neural Networks (NNs) are widely used, and prior work mainly focuses

1 May 2026

AI Value Capture - The Shift To Model Labs

HardwareDGX agent

This article examines how value in the AI industry is shifting away from traditional chip manufacturers toward model development labs and companies that build large language models. It likely analyzes

All the Doomers and hawks are lining up behind this distillation 'attack' farce because they want to see open source banned. It's really as …

HardwareDGX agent

All the Doomers and hawks are lining up behind this distillation 'attack' farce because they want to see open source banned. It's really as simple as that. They want to take away your right to choose,

Efficient Training on Multiple Consumer GPUs with RoundPipe

HardwareDGX agent

arXiv:2604.27085v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on consumer-grade GPUs is highly cost-effective, yet constrained by limited GPU memory and slow PCIe intercon

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

HardwareDGX agent

arXiv:2604.26968v1 Announce Type: cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving. Current

30 Apr 2026

AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices

HardwareDGX agent

arXiv:2604.25326v2 Announce Type: replace-cross Abstract: Speculative decoding enhances the inference efficiency of large language models (LLMs) by generating drafts using a small draft language model

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving

HardwareDGX agent

arXiv:2604.26103v1 Announce Type: cross Abstract: All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneou

Speed Up Unreal Engine NNE Inference with NVIDIA TensorRT for RTX Runtime

HardwareDGX agent

The TensorRT for RTX plugin provides a runtime for Unreal Engine's Neural Network Engine (NNE), enabling efficient deployment of AI models directly within real-time applications. Developers can see 1.

29 Apr 2026

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

Local AiDGX agent

arXiv:2603.15954v2 Announce Type: replace Abstract: Real-time AI experiences call for on-device large language models (OD-LLMs) optimized for efficient deployment on resource-constrained hardware. The

28 Apr 2026

PointTransformerX:Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms

HardwareDGX agent

arXiv:2604.24169v1 Announce Type: new Abstract: 3D point cloud perception remains tightly coupled to custom CUDA operators for spatial operations, limiting portability and efficiency on non-NVIDIA, AM

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

HardwareDGX agent

arXiv:2505.02922v3 Announce Type: replace Abstract: Recent large language models (LLMs) are rapidly extending their context windows, yet inference throughput lags due to increasing GPU memory and band

24 Apr 2026

Bridging the Training-Deployment Gap: Gated Encoding and Multi-Scale Refinement for Efficient Quantization-Aware Image Enhancement

ResearchDGX agent

arXiv:2604.21743v1 Announce Type: new Abstract: Image enhancement models for mobile devices often struggle to balance high output quality with the fast processing speeds required by mobile hardware. W

Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)

HardwareDGX agent

arXiv:2604.16395v2 Announce Type: replace-cross Abstract: Context retrieval systems for LLM inference face a critical challenge: high retrieval latency creates a fundamental tension between waiting fo

23 Apr 2026

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

HardwareDGX agent

arXiv:2412.03594v3 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly play an important role in a wide range of information processing and management tasks in industry. M

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because t…

Model ReleasesDGX agent

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because the headlines are doing you a disservice: Elon Musk got on th

← Previous
1…7891011…75
Next →