AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,730 results
3 Aug 2026

Studying quantization trade-offs for efficient inference deployment in machine translation

HardwareDGX agent

arXiv:2607.29397v1 Announce Type: new Abstract: Deploying large language models in realistic server environments poses challenges, as the system needs to provide high-quality responses with low latenc

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent…

HardwareDGX agent

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent.space/p/inference-eng @Baseten @philipkiely and @waterloo_i

The math still ain’t mathing.

HardwareDGX agent

The math still ain’t mathing. Recently we estimated global (ex China) AI revenues of around 200bn annualised, based on four different sources. Just come across a fifth, based on Nvidia inference sales


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

HardwareDGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

Topology-Aware Data Movement for Disaggregated GPU Inference

HardwareDGX agent

arXiv:2607.28633v1 Announce Type: cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate

2 Aug 2026

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and sever…

HardwareDGX agent

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of opt

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nv…

HardwareDGX agent

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nvidia’s version of a Chinese open model to defend itself agai

I pushed Kimi K3 onto one CPU with 8 GB of RAM

HardwareDGX agent

I deployed K3 on 32 H100s at work a couple of weeks ago and then got annoyed that there was no way to poke at it on my own machine. So I wrote an inference engine for it in C99. Nothing clever going o

The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were …

HardwareDGX agent

The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed door…

HardwareDGX agent

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed doors in just a few organizations doesnt work. What worked in th

1 Aug 2026

Banning open source model, will 'remove capabilities for the defenders' When OpenAI's models breached Huggingface, a Chinese open model is w…

HardwareDGX agent

Banning open source model, will 'remove capabilities for the defenders' When OpenAI's models breached Huggingface, a Chinese open model is what cleaned up the mess. Anthropic's Fable 5 refused, so Hug

Vivid depiction of a dark scenario for Nvidia. “Indeed, CEO Jensen Huang has been so aggressive in funding the AI ecosystem that the company…

HardwareDGX agent

Vivid depiction of a dark scenario for Nvidia. “Indeed, CEO Jensen Huang has been so aggressive in funding the AI ecosystem that the company is functioning almost like a bank; with annual free cash fl

31 Jul 2026

A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes

HardwareDGX agent

arXiv:2607.28047v1 Announce Type: cross Abstract: Time-varying implicit neural representations (INRs) provide a compact representation of scientific volumes and, for modalities such as dynamic X-ray c

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

HardwareDGX agent

arXiv:2607.26661v1 Announce Type: new Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large langu

Autoscaling endpoints for LLM inference

HardwareDGX agent

GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on d

Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

HardwareDGX agent

The article shows that dense‑attention performance in long‑context inference is governed by group size (query heads per KV head), head dimension, and sequence length, with prefill being compute‑bound

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting

HardwareDGX agent

arXiv:2607.27945v1 Announce Type: cross Abstract: Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence learning, nonlinear recurrent updat

FARI: Robust One-Step Inversion for Watermarking in Diffusion Models

HardwareDGX agent

arXiv:2607.26723v1 Announce Type: cross Abstract: Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that i

FasTac: A Curved Multispectral Vision-Based Tactile Sensor for High-Speed High-Precision 3D Shape and Force Perception

HardwareDGX agent

arXiv:2607.28416v1 Announce Type: new Abstract: Curved tactile fingertips for dexterous manipulation must resolve fine contact geometry, distinguish normal and tangential loads, and capture transient

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents

HardwareDGX agent

arXiv:2607.26076v1 Announce Type: cross Abstract: Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse c

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

HardwareDGX agent

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a co

May have found the highest and best use case of Flux 3 - generating GPU ASMR ✨ For everyone who has been asking for access, it’s available N…

HardwareDGX agent

Justine Moore announced on July 31, 2026 that Flux 3’s latest iteration excels at generating GPU‑based ASMR content. She confirmed that this capability is now available in an early preview on the Nous

Nscale buys AI infrastructure optimization startup Anyscale for reported $1.65B

HardwareDGX agent

Data center builder Nscale Global Holdings Ltd. today announced plans to acquire Anyscale Inc., a venture-backed provider of artificial intelligence software. The terms of the deal were not disclosed.

NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek

HardwareDGX agent

NVIDIA Video Codec SDK 13.1 adds AV1 hierarchical reference mode supporting up to 31 B‑frames and efficient iterative tuning that delivers significant bitrate savings in CQ and VBR modes. It enhances

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

HardwareDGX agent

arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches prima

OPENAI IS GOING TO TAKE THIS ENTIRE MARKET DOWN WITH IT And you don't have to own a single share to get hurt. What I'm about to explain shou…

HardwareDGX agent

OPENAI IS GOING TO TAKE THIS ENTIRE MARKET DOWN WITH IT And you don't have to own a single share to get hurt. What I'm about to explain should worry anybody who thinks they're diversified: OpenAI is a

ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate

HardwareDGX agent

arXiv:2607.28144v1 Announce Type: cross Abstract: We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The e

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

HardwareDGX agent

arXiv:2607.28627v1 Announce Type: new Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

HardwareDGX agent

arXiv:2607.27744v1 Announce Type: new Abstract: Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far sy

Security at the foundation requires openness at the foundation. As open-weight models become critical infrastructure for the next generation…

HardwareDGX agent

Security at the foundation requires openness at the foundation. As open-weight models become critical infrastructure for the next generation of software, the systems around them must be transparent, i

Simulation of Surgical Suturing Using Position-Based Dynamics and the Material Point Method for Robot Reinforcement Learning

HardwareDGX agent

arXiv:2607.27494v1 Announce Type: new Abstract: Recent advances in robotics research have created a strong demand for high-performance simulators. Surgical robotics simulation faces unique challenges

Sources: Moonshot has a computing power agreement with Alibaba for the use of ~20K Nvidia chips; some say the deal is for H200 chips, which Alibaba denies (Mackenzie Hawkins/Bloomberg)

HardwareDGX agent

Mackenzie Hawkins / Bloomberg: Sources: Moonshot has a computing power agreement with Alibaba for the use of ~20K Nvidia chips; some say the deal is for H200 chips, which Alibaba denies — Chinese AI c

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

HardwareDGX agent

arXiv:2607.28032v1 Announce Type: new Abstract: Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian

The first thing we learned building autoscaling for dedicated inference is that CPU-style metrics don't tell the whole picture. A GPU can re…

HardwareDGX agent

The first thing we learned building autoscaling for dedicated inference is that CPU-style metrics don't tell the whole picture. A GPU can read 60% busy while the engine's queue is already backing up,

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimo…

HardwareDGX agent

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https://

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

HardwareDGX agent

arXiv:2607.28263v1 Announce Type: new Abstract: Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for pre

30 Jul 2026

BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens

HardwareDGX agent

arXiv:2607.26829v1 Announce Type: new Abstract: Many high-performing volumetric segmentation models maintain dense multi-scale feature maps, leading to high activation memory and inference cost. We pr

Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW

HardwareDGX agent

Back to school means balancing assignments, deadlines and downtime. GeForce NOW makes it easy to have it all. With cloud gaming, everyday laptops used for class can also become GeForce RTX-powered gam

Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI re…

HardwareDGX agent

Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, dec

Can an open-source model perform like a foundation model? @Osmosis_AI is betting yes, using reinforcement learning and the dedicated @ycombi…

HardwareDGX agent

Osmosis_AI claims that an open‑source model can rival a foundation model by leveraging reinforcement learning techniques. To demonstrate this, they will use the Y Combinator‑dedicated GPU cluster on T

Four Ways to Deploy More Secure AI Agents

HardwareDGX agent

NVIDIA’s AI Red Team found common failure modes in enterprise AI agents: weak access controls, unrestricted code execution via tools, unprotected network egress, and exposure of plaintext secrets. The

Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation

HardwareDGX agent

arXiv:2607.26646v1 Announce Type: new Abstract: We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single 360^irc panorama, without per-scene optimization or mu

I built ganfs: A Python package that uses GANs to automate feature selection for high-dimensional datasets. (No domain expert required) [P] [R]

HardwareDGX agent

Hey everyone, I recently open-sourced a new Python package called ganfs (Generative Adversarial Network Feature Selection), and I wanted to share it with the community. The Problem: Selecting the best

Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment

HardwareDGX agent

arXiv:2607.26238v1 Announce Type: new Abstract: We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. Using DINOv2-L (304M parame

NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and Education

HardwareDGX agent

arXiv:2607.26855v1 Announce Type: new Abstract: Many scientific fields rely on standard benchmarks and shared platforms to improve review and reproducibility, but autonomous systems research still lac

NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure

HardwareDGX agent

NVIDIA’s Exemplar Cloud study shows that identical H100‑based clusters can yield 8–12 % lower training throughput when configuration gaps—at the kernel, hypervisor, BIOS or NCCL levels—prevent reachin

PolygMap: A Perceptive Locomotion Framework for Humanoid Robot Stair Climbing

HardwareDGX agent

arXiv:2510.12346v2 Announce Type: replace Abstract: Recently, biped robot walking technology has been significantly developed, mainly in the context of a bland walking scheme. To emulate human walking

Protopia and Rafay deliver multi-tenancy for shared GPU AI factories

HardwareDGX agent

Enterprise AI infrastructure providers are turning to multi-tenancy paired with upstream data protection to convert idle GPU capacity into secure, token-metered services that enterprises will actually

Q&A with CuspAI's Max Welling on its AI Materials Foundry, partnerships with Nvidia and others, Geoff Hinton and Yann LeCun joining its advisory board, and more (John Thornhill/Financial Times)

HardwareDGX agent

John Thornhill / Financial Times: Q&A with CuspAI's Max Welling on its AI Materials Foundry, partnerships with Nvidia and others, Geoff Hinton and Yann LeCun joining its advisory board, and more — The

Run High-Performance Core Math at Scale with NVIDIA nvmath-python

HardwareDGX agent

NVIDIA **nvmath‑python v1.0** is a Python library that wraps CUDA‑X and NVPL math libraries (cuFFT, cuBLASLt, cuDSS, cuSPARSE, cuTENSOR, cuBLASMp) to give NumPy, CuPy, and PyTorch users GPU‑accelerate

unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose?

HardwareDGX agent

Are there any benchmarks on these 4 bit quants, like how Artificial Analysis runs a slew of various benchmarks? If not, how can I run one (5x over for consistency) on them? I'm also very interested in

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

HardwareDGX agent

arXiv:2607.26694v1 Announce Type: new Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation,

What to expect during Black Hat USA: Join theCUBE Aug. 5-6

HardwareDGX agent

Artificial intelligence has propelled the cybersecurity world into a new phase, driven by startlingly advanced autonomous attacks and an urgent need to adopt technology to defend against them. This we

29 Jul 2026

agreed. which is part of why coating the world in data centers is a profound mistake.

HardwareDGX agent

agreed. which is part of why coating the world in data centers is a profound mistake. AI will get so ridiculously efficient that we will look back at GPU clusters the way we now look at these first ro

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

HardwareDGX agent

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar

HardwareDGX agent

arXiv:2607.26016v1 Announce Type: cross Abstract: Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic a

Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging

HardwareDGX agent

arXiv:2607.25967v1 Announce Type: new Abstract: Singular Value Decomposition (SVD) underlies matrix factorisation tasks across computational imaging, with medical applications increasingly demanding r

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a …

HardwareDGX agent

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a lot with agent reliability. More below: Agent development to

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

HardwareDGX agent

arXiv:2505.14366v2 Announce Type: replace Abstract: We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embod

28 Jul 2026

A didactical-driven teacher assistant for a dimensional modeling course

HardwareDGX agent

arXiv:2607.22598v1 Announce Type: cross Abstract: Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisi

← Previous
12345…29
Next →