AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,730 results
12 Aug 2026

CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

HardwareDGX agent

arXiv:2608.10506v1 Announce Type: cross Abstract: Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resourc

EvoMem: Memory-Augmented Evolution for Code Optimization

HardwareDGX agent

arXiv:2608.10795v1 Announce Type: new Abstract: Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may tran

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

HardwareDGX agent

arXiv:2608.10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

HardwareDGX agent

arXiv:2608.10317v1 Announce Type: new Abstract: We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection.

Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4

HardwareDGX agent

arXiv:2608.10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads wit

HyWA: Architecture-Preserving Personalized Voice Activity Detection for Full-Duplex Voice Assistants

HardwareDGX agent

arXiv:2510.12947v3 Announce Type: replace-cross Abstract: Voice activity detection (VAD) serves as an early gate in voice-assistant pipelines for smart devices. Because conventional VADs respond to sp

IBM inks $240M infrastructure deal with AI-optimized cloud operator Together AI

HardwareDGX agent

IBM Corp. will provide infrastructure to artificial intelligence startup Together AI Inc. as part of a 240 million deal announced today. The partnership comes a few weeks after the latter company rais

Is the future of AI selling hardware for Open Source/Models?

HardwareDGX agent

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c

Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures

HardwareDGX agent

arXiv:2608.10323v1 Announce Type: new Abstract: Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation. We present Neuroevolution Arena,

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

HardwareDGX agent

We announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobilize over $500 billion of third-party capit

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

HardwareDGX agent

arXiv:2608.10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external e

Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine

HardwareDGX agent

Running large language model inference at scale forces a KV cache trade-off: oversized GPU instances or slow time-to-first-token. This post builds a tiered KV cache on Amazon SageMaker HyperPod that e

TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification

HardwareDGX agent

arXiv:2504.11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual su

What do you guys do for GPU Kernels?

HardwareDGX agent

I'm trying to figure out GPU Kernel optimization on older hardware like SM80(ampere) . Is there tools you guys use? Or frameworks? Im waiting for this framework https://www.reddit.com/r/LocalLLaMA/com

11 Aug 2026

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation

HardwareDGX agent

arXiv:2608.08469v1 Announce Type: new Abstract: Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duple

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

HardwareDGX agent

arXiv:2608.08256v1 Announce Type: new Abstract: Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but

Beyond Isotropic Assumptions: Continuity-Constrained Segmentation and GPU Morphometry for Nanoscale GBM Analysis

HardwareDGX agent

arXiv:2608.07575v1 Announce Type: new Abstract: Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic:

Big Tech's AI boom echoes the 1870s railroad buildout, and Nvidia shifting risk to institutional capital may expose investors if AI revenues fail to materialize (Ben Thompson/Stratechery)

HardwareDGX agent

Ben Thompson / Stratechery: Big Tech's AI boom echoes the 1870s railroad buildout, and Nvidia shifting risk to institutional capital may expose investors if AI revenues fail to materialize — On Januar

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

HardwareDGX agent

arXiv:2608.09444v1 Announce Type: cross Abstract: A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the mode

Design Space of Self--Consistent Electrostatic Machine Learning Interatomic Potentials

HardwareDGX agent

arXiv:2603.14700v2 Announce Type: replace-cross Abstract: Machine learning interatomic potentials (MLIPs) have become widely used tools in atomistic simulations. For much of the history of this field,

Differentiate the Solver, Not the Equation: Reverse-Sweep Adjoints for Block Implicit Simulation

HardwareDGX agent

arXiv:2608.08559v1 Announce Type: cross Abstract: Differentiable simulation is a key component in learning, control, and inverse problems, where gradients through nonlinear implicit solvers are requir

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

HardwareDGX agent

arXiv:2608.07964v1 Announce Type: cross Abstract: Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing distributions

Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence

HardwareDGX agent

arXiv:2608.08761v1 Announce Type: cross Abstract: In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturin

ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB Viewpoints

HardwareDGX agent

arXiv:2608.08531v1 Announce Type: new Abstract: Deep learning-driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D

EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac Sim

HardwareDGX agent

arXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

HardwareDGX agent

arXiv:2607.18171v2 Announce Type: replace Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose effici

From Approachability Residuals to Anytime-Valid Evidence: The Online Convex Geometry of Testing by Betting

HardwareDGX agent

arXiv:2608.09450v1 Announce Type: new Abstract: Betting-based sequential tests and Blackwell approachability are linked by a rate-explicit reduction through support-function residuals. For a compact c

Holder Signed Distance: A Differentiable, Signed, Parallelizable Metric for Robotics

HardwareDGX agent

arXiv:2608.07707v1 Announce Type: new Abstract: Computing distances between sets is essential in robotic motion planning and control, where differentiable gradients enable real-time optimization. The

IBM and Together AI sign a $240M, multi-year deal to build an AI inference cluster on IBM Cloud, using Nvidia's HGX B300 systems, to support open-source models (Anhata Rooprai/Reuters)

HardwareDGX agent

Anhata Rooprai / Reuters: IBM and Together AI sign a 240M, multi-year deal to build an AI inference cluster on IBM Cloud, using Nvidia's HGX B300 systems, to support open-source models — IBM (IBM.N) a

Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality

HardwareDGX agent

arXiv:2512.20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited paralle

More on our partnership: https://newsroom.ibm.com/2026-08-11-IBM-and-Together-AI-Sign-Multi-Year-Agreement-to-Scale-Open-Source-AI-Inference…

HardwareDGX agent

In August 2026, Together AI entered into a multi‑year partnership with IBM and NVIDIA to deliver enterprise‑grade open‑source AI inference on IBM Cloud. The collaboration deploys a dedicated NVIDIA B3

Multi-tier storage rewrites the economics of AI inference

HardwareDGX agent

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance. These architectures combine fl

🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sa…

HardwareDGX agent

🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sachs, and KKR, to raise over $500 billion so its own customers

NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation

HardwareDGX agent

NVIDIA JetPack 7.2.1 adds agentic video skills with the unified jetson‑videosdk, allowing programmable, device-aware video workflows that link developer intent to live device discovery and performance

Nvidia Nemo Switchyard

HardwareDGX agent

https://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i

Nvidia taps Wall Street for a half-trillion dollars to fuel global AI infrastructure buildout

HardwareDGX agent

Some of Wall Street’s biggest financial firms are partnering with Nvidia Corp. to pour a half-trillion dollars of funding into the artificial intelligence industry’s massive infrastructure buildout. N

Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD

HardwareDGX agent

River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised 1.1 billion in early-stage funding. The company stated in today’s announcement that it

Physics-Informed Condition Monitoring of SiC Power Modules

HardwareDGX agent

arXiv:2608.08363v1 Announce Type: cross Abstract: Silicon carbide (SiC) power modules are increasingly deployed in automotive traction inverters, where condition monitoring is essential to prevent in-

Real-time physics inversion for retrieval of sub-pixel wildfire temperatures from VSWIR imaging spectroscopy

HardwareDGX agent

arXiv:2608.07580v1 Announce Type: new Abstract: In this work, we present a wildfire temperature retrieval framework for VSWIR imaging spectroscopy data, employed on data from NASA's Airborne Visible I

Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard

HardwareDGX agent

NVIDIA NeMo Switchyard is a routing platform that directs AI agent workloads to the most suitable specialized or frontier model for each step of a task, balancing performance, cost, and latency. It of

StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning

HardwareDGX agent

arXiv:2603.02637v2 Announce Type: replace-cross Abstract: Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to depende

SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization

HardwareDGX agent

arXiv:2608.09160v1 Announce Type: new Abstract: Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectr…

HardwareDGX agent

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectrum-X networking. First of its kind on IBM Cloud, powered by

verdi: retrieval is not transfer for continual world model optimization

HardwareDGX agent

arXiv:2608.09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model t

What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload

HardwareDGX agent

arXiv:2608.08287v1 Announce Type: new Abstract: GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement

Why Scaling AI Compute Performance Requires a New Power Architecture

HardwareDGX agent

Every new generation of accelerated computing demands more from the infrastructure underneath it — more compute performance, higher rack density and more efficient, scalable power distribution. The bo

10 Aug 2026

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default …

HardwareDGX agent

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default because it sees queue pressure before latency degrades. We t

Beyond Visibility: Real-Time Surface Accessibility Fields from Sparse LiDAR

HardwareDGX agent

arXiv:2608.06412v1 Announce Type: cross Abstract: Understanding which surfaces in a scene are physically accessible to a given tool is fundamental for robotic interaction, yet 3D perception systems ty

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights

HardwareDGX agent

arXiv:2608.06763v1 Announce Type: new Abstract: Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU

Fast LapSum: Exact Differentiable Top-k at Million Scale

HardwareDGX agent

arXiv:2608.06912v1 Announce Type: new Abstract: The top-k operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and atten

If Anthropic begin shipping chips, will NVIDIA begin shipping frontier models? Who will win?

HardwareDGX agent

On August 10 2026 at 1:39 AM UTC, user Itamar Friedman (@itamar_mar) posted a short tweet asking whether Anthropic’s potential launch of its own chips would prompt NVIDIA to release frontier AI models

Jensen Huang just told the incredible story of how Elon Musk became NVIDIA’s first customer for its AI supercomputer - when literally nobody…

HardwareDGX agent

Jensen Huang just told the incredible story of how Elon Musk became NVIDIA’s first customer for its AI supercomputer - when literally nobody else wanted it “When I announced this thing, nobody wanted

Meta says Muse Glimmer has 30B parameters and is 'small enough' to need only one GPU; Meta plans a $1B fund to invest in US communities near its data centers (Vlad Savov/Bloomberg)

HardwareDGX agent

Vlad Savov / Bloomberg: Meta says Muse Glimmer has 30B parameters and is “small enough” to need only one GPU; Meta plans a $1B fund to invest in US communities near its data centers — Meta Platforms I

Nvidia partners with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR on a $500B funding package for AI infrastructure development (Financial Times)

HardwareDGX agent

Financial Times: Nvidia partners with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR on a $500B funding package for AI infrastructure development — Apollo, Blackstone and Goldman Sa

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

HardwareDGX agent

arXiv:2608.07078v1 Announce Type: cross Abstract: Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among exi

Sources: AI cloud computing provider Lambda is selling a $917M leveraged loan to finance the purchase of GPUs as part of a contract with Nvidia (Bloomberg)

HardwareDGX agent

Bloomberg: Sources: AI cloud computing provider Lambda is selling a $917M leveraged loan to finance the purchase of GPUs as part of a contract with Nvidia — Lambda Inc., an AI cloud-computing provider

Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

HardwareDGX agent

arXiv:2608.06959v1 Announce Type: new Abstract: Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downl

Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX

HardwareDGX agent

The TileRT InferenceX article (Aug 10 2026) examines whether the TileRT software stack on NVIDIA GPUs can compete with dedicated inference systems such as Cerebras, Groq LPUs and SambaNova for ultra‑h

Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Bat…

HardwareDGX agent

Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, High throughput prefill eng

9 Aug 2026

300b on 32gb MoE-streaming findings + optimisations

HardwareDGX agent

The past week I've been running DSv4 inference on my laptop by keeping everything RAM-resident except the MXFP4-experts (since expert pool is ~147GB and won't fit) TL;DR - read speed is the limiter mo

← Previous
123…29
Next →