AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,450 results
11 Aug 2026

I ran Muse Glimmer @ 1M context - All tests passed.

Model ReleasesDGX agent

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

IBM and Together AI sign a $240M, multi-year deal to build an AI inference cluster on IBM Cloud, using Nvidia's HGX B300 systems, to support open-source models (Anhata Rooprai/Reuters)

HardwareDGX agent

Anhata Rooprai / Reuters: IBM and Together AI sign a 240M, multi-year deal to build an AI inference cluster on IBM Cloud, using Nvidia's HGX B300 systems, to support open-source models — IBM (IBM.N) a

Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality

HardwareDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited paralle

More on our partnership: https://newsroom.ibm.com/2026-08-11-IBM-and-Together-AI-Sign-Multi-Year-Agreement-to-Scale-Open-Source-AI-Inference…

HardwareDGX agent

In August 2026, Together AI entered into a multi‑year partnership with IBM and NVIDIA to deliver enterprise‑grade open‑source AI inference on IBM Cloud. The collaboration deploys a dedicated NVIDIA B3

Multi-tier storage rewrites the economics of AI inference

HardwareDGX agent

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance. These architectures combine fl

NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation

HardwareDGX agent

NVIDIA JetPack 7.2.1 adds agentic video skills with the unified jetson‑videosdk, allowing programmable, device-aware video workflows that link developer intent to live device discovery and performance

Nvidia Nemo Switchyard

HardwareDGX agent

https://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i

Nvidia taps Wall Street for a half-trillion dollars to fuel global AI infrastructure buildout

HardwareDGX agent

Some of Wall Street’s biggest financial firms are partnering with Nvidia Corp. to pour a half-trillion dollars of funding into the artificial intelligence industry’s massive infrastructure buildout. N

Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD

HardwareDGX agent

River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised 1.1 billion in early-stage funding. The company stated in today’s announcement that it

Physics-Informed Condition Monitoring of SiC Power Modules

HardwareDGX agent

arXiv:2608.08363v1 Announce Type: cross Abstract: Silicon carbide (SiC) power modules are increasingly deployed in automotive traction inverters, where condition monitoring is essential to prevent in-

Real-time physics inversion for retrieval of sub-pixel wildfire temperatures from VSWIR imaging spectroscopy

HardwareDGX agent

arXiv:2608.07580v1 Announce Type: new Abstract: In this work, we present a wildfire temperature retrieval framework for VSWIR imaging spectroscopy data, employed on data from NASA's Airborne Visible I

Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard

HardwareDGX agent

NVIDIA NeMo Switchyard is a routing platform that directs AI agent workloads to the most suitable specialized or frontier model for each step of a task, balancing performance, cost, and latency. It of

StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning

HardwareDGX agent

arXiv:2603.02637v2 Announce Type: replace-cross Abstract: Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to depende

SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization

HardwareDGX agent

arXiv:2608.09160v1 Announce Type: new Abstract: Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism

Thought-Level Beam Search for Reasoning

ResearchDGX agent

arXiv:2608.08020v1 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shift

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectr…

HardwareDGX agent

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectrum-X networking. First of its kind on IBM Cloud, powered by

verdi: retrieval is not transfer for continual world model optimization

HardwareDGX agent

arXiv:2608.09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model t

Why Scaling AI Compute Performance Requires a New Power Architecture

HardwareDGX agent

Every new generation of accelerated computing demands more from the infrastructure underneath it — more compute performance, higher rack density and more efficient, scalable power distribution. The bo

10 Aug 2026

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default …

HardwareDGX agent

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default because it sees queue pressure before latency degrades. We t

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights

HardwareDGX agent

arXiv:2608.06763v1 Announce Type: new Abstract: Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU

Fast LapSum: Exact Differentiable Top-k at Million Scale

HardwareDGX agent

arXiv:2608.06912v1 Announce Type: new Abstract: The top-k operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and atten

If Anthropic begin shipping chips, will NVIDIA begin shipping frontier models? Who will win?

HardwareDGX agent

On August 10 2026 at 1:39 AM UTC, user Itamar Friedman (@itamar_mar) posted a short tweet asking whether Anthropic’s potential launch of its own chips would prompt NVIDIA to release frontier AI models

Jensen Huang just told the incredible story of how Elon Musk became NVIDIA’s first customer for its AI supercomputer - when literally nobody…

HardwareDGX agent

Jensen Huang just told the incredible story of how Elon Musk became NVIDIA’s first customer for its AI supercomputer - when literally nobody else wanted it “When I announced this thing, nobody wanted

Meta says Muse Glimmer has 30B parameters and is 'small enough' to need only one GPU; Meta plans a $1B fund to invest in US communities near its data centers (Vlad Savov/Bloomberg)

HardwareDGX agent

Vlad Savov / Bloomberg: Meta says Muse Glimmer has 30B parameters and is “small enough” to need only one GPU; Meta plans a $1B fund to invest in US communities near its data centers — Meta Platforms I

Nvidia partners with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR on a $500B funding package for AI infrastructure development (Financial Times)

HardwareDGX agent

Financial Times: Nvidia partners with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR on a $500B funding package for AI infrastructure development — Apollo, Blackstone and Goldman Sa

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

HardwareDGX agent

arXiv:2608.07078v1 Announce Type: cross Abstract: Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among exi

Sources: AI cloud computing provider Lambda is selling a $917M leveraged loan to finance the purchase of GPUs as part of a contract with Nvidia (Bloomberg)

HardwareDGX agent

Bloomberg: Sources: AI cloud computing provider Lambda is selling a $917M leveraged loan to finance the purchase of GPUs as part of a contract with Nvidia — Lambda Inc., an AI cloud-computing provider

Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

HardwareDGX agent

arXiv:2608.06959v1 Announce Type: new Abstract: Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downl

Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX

HardwareDGX agent

The TileRT InferenceX article (Aug 10 2026) examines whether the TileRT software stack on NVIDIA GPUs can compete with dedicated inference systems such as Cerebras, Groq LPUs and SambaNova for ultra‑h

9 Aug 2026

300b on 32gb MoE-streaming findings + optimisations

HardwareDGX agent

The past week I've been running DSv4 inference on my laptop by keeping everything RAM-resident except the MXFP4-experts (since expert pool is ~147GB and won't fit) TL;DR - read speed is the limiter mo

I Turned My Underused Gaming Laptop Into a Local AI Workstation

Local AiDGX agent

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission

Open Model: Google Weather Next 2

HardwareDGX agent

I am not a meteorologist, but I just read a very interesting article: https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day/ In a paper published on Thursda

7 Aug 2026

CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

HardwareDGX agent

arXiv:2512.02551v4 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimi

Echo Dot 2 can run 28M LLM at decent speed

Model ReleasesDGX agent

Code and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running

IcFuzz: Fuzzing Isaac Sim with Semantic Stage Guidance and Multi-level Mutation

HardwareDGX agent

arXiv:2608.06088v1 Announce Type: new Abstract: Robotics simulators serve as a foundational infrastructure for embodied AI, facilitating safe and scalable robotic system development. NVIDIA Isaac Sim

LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

Model ReleasesDGX agent

LabyrinthBench measures the thing that actually kills long agent runs — whether a model can still use what it learned twenty turns ago — deterministically, with no LLM judge, on your own hardware, wit

Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation

HardwareDGX agent

arXiv:2608.05819v1 Announce Type: new Abstract: Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network co

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

HardwareDGX agent

arXiv:2608.06146v1 Announce Type: new Abstract: End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formula

Sources: Nvidia agrees to invest 2B in Lancium, the power infrastructure developer of the Stargate campus in Texas, plus 1B more if it hits certain thresholds (The Information)

HardwareDGX agent

The Information: Sources: Nvidia agrees to invest 2B in Lancium, the power infrastructure developer of the Stargate campus in Texas, plus 1B more if it hits certain thresholds — Nvidia has agreed to i

Sources: the US Commerce Department's BIS is reviewing how Chinese AI companies access Nvidia chips overseas, including by legally renting foreign data centers (Mackenzie Hawkins/Bloomberg)

HardwareDGX agent

Mackenzie Hawkins / Bloomberg: Sources: the US Commerce Department's BIS is reviewing how Chinese AI companies access Nvidia chips overseas, including by legally renting foreign data centers — A key U

6 Aug 2026

AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance

HardwareDGX agent

arXiv:2608.05109v1 Announce Type: cross Abstract: Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional frin

AMD acquires Taalas to hardwire AI models into silicon

HardwareDGX agent

Advanced Micro Devices Inc. said today it has agreed to buy Taalas Inc., a Toronto startup that hardwires artificial intelligence models directly into silicon, in a deal that pushes the chipmaker deep

AMD acquires Toronto-based Taalas, which integrates model weights directly into silicon to boost inference performance, for an undisclosed sum (Tobias Mann/The Register)

HardwareDGX agent

Tobias Mann / The Register: AMD acquires Toronto-based Taalas, which integrates model weights directly into silicon to boost inference performance, for an undisclosed sum — Early tech demos show model

An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells

HardwareDGX agent

arXiv:2608.04041v1 Announce Type: new Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classific

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

HardwareDGX agent

arXiv:2608.04956v1 Announce Type: new Abstract: Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate op

Docs: US data labeling companies, like Surge AI and Mercor, that sell training datasets to US AI labs and the government are also selling them to Chinese labs (Anna Tong/Forbes)

HardwareDGX agent

Anna Tong / Forbes: Docs: US data labeling companies, like Surge AI and Mercor, that sell training datasets to US AI labs and the government are also selling them to Chinese labs — The same Silicon Va

GASP: GPU-Accelerated Safe Planner for Real-Time Collision-Aware Motion Generation with Latent Trajectory Sampling

HardwareDGX agent

arXiv:2608.04612v1 Announce Type: new Abstract: We present GASP, a GPU-Accelerated Safe Planner for real-time, collision-aware joint-space motion generation in known environments. GASP combines a clam

Gaussian-LIC2: LiDAR-Inertial-Camera Gaussian Splatting SLAM

HardwareDGX agent

arXiv:2507.04004v3 Announce Type: replace Abstract: This paper presents the first photo-realistic LiDAR-Inertial-Camera Gaussian Splatting SLAM system that simultaneously addresses visual quality, geo

Hugging Face Storage Buckets are now on http://Vast.ai Connect your HF Storage Bucket as a Cloud Connection in your Vast settings, and every…

HardwareDGX agent

Hugging Face Storage Buckets are now on http://Vast.ai Connect your HF Storage Bucket as a Cloud Connection in your Vast settings, and every GPU instance you rent can pull datasets and checkpoints str

Into the Omniverse: How Open World Models Push the Frontier of Physical AI

HardwareDGX agent

In July, NVIDIA joined more than 200 companies and organizations in signing “Open Weights and American AI Leadership,” an open letter arguing that AI leadership will be measured not by any single fron

Jony Ive’s first OpenAI gadget is reportedly a hockey puck-sized smart speaker

IndustryDGX agent

The AI device OpenAI is developing with former Apple designer Jony Ive is 'essentially a smart speaker without a display' that's battery-powered, doughnut-shaped and roughly the size of a hockey puck,

Ooredoo, Nvidia, Nokia, and Indosat launch Zankore, Indonesia's first dedicated AI compute and neocloud platform; Ooredoo, which owns a 49% stake, commits $800M (Melissa Hancock/Fortune)

HardwareDGX agent

Melissa Hancock / Fortune: Ooredoo, Nvidia, Nokia, and Indosat launch Zankore, Indonesia's first dedicated AI compute and neocloud platform; Ooredoo, which owns a 49% stake, commits $800M — Ooredoo Gr

RORA: Realistic Object Reconstruction with Articulation

HardwareDGX agent

arXiv:2608.04842v1 Announce Type: new Abstract: Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effe

StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning

HardwareDGX agent

arXiv:2608.04811v1 Announce Type: new Abstract: Structured pruning enhances the efficiency of deep neural networks (DNNs) by eliminating groups of parameters during inference. Previous methods mostly

Sydney-based AI data center company Firmus raised 2B in funding from Coatue, Nvidia, and others at a 10.5B post-money valuation, up from $5.5B in April (Nichiket Sunil/Reuters)

HardwareDGX agent

Nichiket Sunil / Reuters: Sydney-based AI data center company Firmus raised 2B in funding from Coatue, Nvidia, and others at a 10.5B post-money valuation, up from 5.5B in April — Firmus said on Friday

5 Aug 2026

Accelerating Dynamic Graph Clustering on GPU Architectures with cuGraph

HardwareDGX agent

arXiv:2608.03695v1 Announce Type: cross Abstract: This work addresses community detection in temporal networks through GPU-accelerated extensions of spectral clustering and modularity-based algorithms

AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding

HardwareDGX agent

arXiv:2608.02989v1 Announce Type: cross Abstract: Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verifi

Compound and Parallel Modes of Tropical Convolutional Neural Networks

HardwareDGX agent

arXiv:2504.06881v2 Announce Type: replace-cross Abstract: Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems, yet their reliance on multiplication-i

Learning and Clustering on Temporal Graphs: Principles, Primitives, and Pooling

HardwareDGX agent

arXiv:2608.03696v1 Announce Type: new Abstract: This work focuses on the problem of learning on temporal graphs, with particular emphasis on the task of clustering: obtaining coarse-grained representa

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

HardwareDGX agent

arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticip

← Previous
1…910111213…75
Next →