AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
4,457 results
5 Aug 2026

Learning and Clustering on Temporal Graphs: Principles, Primitives, and Pooling

HardwareDGX agent

arXiv:2608.03696v1 Announce Type: new Abstract: This work focuses on the problem of learning on temporal graphs, with particular emphasis on the task of clustering: obtaining coarse-grained representa

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

HardwareDGX agent

arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticip

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs

HardwareDGX agent

arXiv:2608.03036v1 Announce Type: cross Abstract: Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Se

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RTX 3090 MiniMax H3 Speed Comparison: FP8 Scaled vs INT8 ConvRot (W8A8)

HardwareDGX agent

Setup: GPU: RTX 3090 24GB RAM: 32GB ComfyUI 0.30.0 PyTorch 2.13.0+cu130 CUDA 13.0 SageAttention enabled (sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64) Spectrum node + Euler

4 Aug 2026

A Fortran General-Purpose Transpiler: Proof of Concept

HardwareDGX agent

arXiv:2608.00130v1 Announce Type: cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an expertise

AI cloud startup Volta Infra raised 300M led by a16z and Altimeter at a 2.4B valuation and says it landed a $10B contract with an unnamed leading AI developer (Dina Bass/Bloomberg)

HardwareDGX agent

Dina Bass / Bloomberg: AI cloud startup Volta Infra raised 300M led by a16z and Altimeter at a 2.4B valuation and says it landed a 10B contract with an unnamed leading AI developer — Volta Infra Holdi

AI compute is going to orbit. 🚀 @SpaceX’s Starmind AI1 satellite compute payload is powered by NVIDIA Vera Rubin NVL72, bringing AI factory…

HardwareDGX agent

AI compute is going to orbit. 🚀 @SpaceX’s Starmind AI1 satellite compute payload is powered by NVIDIA Vera Rubin NVL72, bringing AI factory compute closer to the stars. The next chapter of AI infrastr

As AI Increases Demands on Memory, Storage Steps Up

HardwareDGX agent

Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory. But rising needs aren’t met by simply adding more storage capacity. What

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

HardwareDGX agent

arXiv:2608.01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autore

Choosing the right orchestration layer for your AI use cases

HardwareDGX agent

Compute scarcity is not the only struggle AI teams face. Optimal utilization is also key. Friction also emerges when they outgrow informal coordination methods such as shared spreadsheets, manual SSH

early sparks of rsi? Ali Taha @waterloo_intern explains how http://Z.ai's GLM-5.2 (@Zai_org) profiled its SGLang serving path and rewrote bo…

HardwareDGX agent

early sparks of rsi? Ali Taha @waterloo_intern explains how http://Z.ai's GLM-5.2 (@Zai_org) profiled its SGLang serving path and rewrote bottlenecked GPU kernels, (still lacks reliable judgment) grea

Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super

HardwareDGX agent

NVIDIA Alpamayo 2 Super is a publicly available 34‑billion‑parameter vision–language–action model that merges a 32‑B Cosmos 3 Super Reasoner with a 2‑B action‑expert diffusion network. It produces uni

Has anyone been working on a solid setup for DSV4F on x2+ R9700s?

HardwareDGX agent

I'm hoping that one of you guys has been working on an inference engine or has somehow found improvements to running DSV4F on RDNA4 multi-GPU setups. I am currently building a custom inference engine

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference

HardwareDGX agent

arXiv:2608.00577v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous

MiniWorld: Democratizing the Training of Video World Models from Scratch

HardwareDGX agent

arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through auto

Nova: An End-to-End MLIR Compiler for Deep Learning

Model ReleasesDGX agent

arXiv:2608.00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physica

NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

HardwareDGX agent

For robotaxis and other autonomous vehicles (AVs), the hardest problems aren’t the everyday scenarios. They’re the rare, complex situations that are difficult to anticipate and train for. Handling the

NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US

HardwareDGX agent

NVIDIA is participating in the U.S. National Science Foundation’s (NSF) State and Regional Artificial Intelligence Infrastructure Hubs program, an effort launching today to expand access to the advanc

Nvidia open sources cuFile API, accelerating GPU read/write capability for high-speed storage

HardwareDGX agent

As artificial intelligence applications become ever hungrier for faster access to data, Nvidia Corp. today announced it is open-sourcing the application programming interface for its powerful cuFile v

Open Secure AI Alliance proposes SAFE guidelines as membership tops 120

HardwareDGX agent

The Open Secure AI Alliance today proposed a set of guidelines for reporting cybersecurity incidents involving artificial intelligence agents, one week after the group was formed. The proposal is call

Partial FC: Training 10 Million Identities on a Single Machine

HardwareDGX agent

arXiv:2010.05222v3 Announce Type: replace Abstract: Training face recognition models with millions of identities is challenging because classifier storage, logit memory, and computation grow linearly

Sources: Anthropic agreed to a $10B deal for computing capacity in Norway from Nvidia-backed AI cloud startup Volta Infra, which says the deal is for six years (Bloomberg)

HardwareDGX agent

Bloomberg: Sources: Anthropic agreed to a 10B deal for computing capacity in Norway from Nvidia-backed AI cloud startup Volta Infra, which says the deal is for six years — Anthropic PBC has struck a 1

Stipple: Real-Time Incremental Gaussian Splatting with Visual-Inertial Tracking

HardwareDGX agent

arXiv:2608.00931v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) provides efficient rendering of photo-realistic scenes, but its heavy preprocessing and training steps make it a poor fit

Toward Geometry-Scalable Whole-Body Touch for Humanoids: A 3D-Printed Conformal EIT Skin

HardwareDGX agent

arXiv:2608.02080v1 Announce Type: new Abstract: Whole-body tactile sensing is a prerequisite for humanoids that operate in contact-rich human environments, but conventional taxel arrays scale poorly w

3 Aug 2026

Event-Based Upper-Body Humanoid Teleoperation Under Challenging Illumination

HardwareDGX agent

arXiv:2607.29227v1 Announce Type: new Abstract: We present a real-time upper-body human-to-humanoid motion imitation framework driven by neuromorphic event-based vision. This work addresses practical

GPU-Accelerated ANNS: Quantized for Speed, Built for Change

HardwareDGX agent

arXiv:2601.07048v5 Announce Type: replace-cross Abstract: Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promisin

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

HardwareDGX agent

arXiv:2607.29158v1 Announce Type: cross Abstract: We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point

London-based AI chip startup Olix raised 312M led by Fundomo, with participation from Arm, at a 3.3B valuation, up from 1B+ after raising 220M in February (Tim Bradshaw/Financial Times)

HardwareDGX agent

Tim Bradshaw / Financial Times: London-based AI chip startup Olix raised 312M led by Fundomo, with participation from Arm, at a 3.3B valuation, up from 1B+ after raising 220M in February — British ent

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage

HardwareDGX agent

NVIDIA’s Vera BlueField‑4 STX Storage Processor combines an 88‑core Armv9.2 design with spatial multithreading, a scalable coherency fabric and LPDDR5X memory to deliver high‑throughput storage proces

Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds

HardwareDGX agent

arXiv:2607.28994v1 Announce Type: cross Abstract: High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit p

Studying quantization trade-offs for efficient inference deployment in machine translation

HardwareDGX agent

arXiv:2607.29397v1 Announce Type: new Abstract: Deploying large language models in realistic server environments poses challenges, as the system needs to provide high-quality responses with low latenc

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent…

HardwareDGX agent

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent.space/p/inference-eng @Baseten @philipkiely and @waterloo_i

The math still ain’t mathing.

HardwareDGX agent

The math still ain’t mathing. Recently we estimated global (ex China) AI revenues of around 200bn annualised, based on four different sources. Just come across a fifth, based on Nvidia inference sales

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

HardwareDGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

2 Aug 2026

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and sever…

HardwareDGX agent

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of opt

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nv…

HardwareDGX agent

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nvidia’s version of a Chinese open model to defend itself agai

I pushed Kimi K3 onto one CPU with 8 GB of RAM

HardwareDGX agent

I deployed K3 on 32 H100s at work a couple of weeks ago and then got annoyed that there was no way to poke at it on my own machine. So I wrote an inference engine for it in C99. Nothing clever going o

The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were …

HardwareDGX agent

The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed door…

HardwareDGX agent

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed doors in just a few organizations doesnt work. What worked in th

1 Aug 2026

Banning open source model, will 'remove capabilities for the defenders' When OpenAI's models breached Huggingface, a Chinese open model is w…

HardwareDGX agent

Banning open source model, will 'remove capabilities for the defenders' When OpenAI's models breached Huggingface, a Chinese open model is what cleaned up the mess. Anthropic's Fable 5 refused, so Hug

31 Jul 2026

A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes

HardwareDGX agent

arXiv:2607.28047v1 Announce Type: cross Abstract: Time-varying implicit neural representations (INRs) provide a compact representation of scientific volumes and, for modalities such as dynamic X-ray c

Autoscaling endpoints for LLM inference

HardwareDGX agent

GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on d

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting

HardwareDGX agent

arXiv:2607.27945v1 Announce Type: cross Abstract: Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence learning, nonlinear recurrent updat

FARI: Robust One-Step Inversion for Watermarking in Diffusion Models

HardwareDGX agent

arXiv:2607.26723v1 Announce Type: cross Abstract: Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that i

Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras

ApplicationsDGX agent

arXiv:2607.28293v1 Announce Type: new Abstract: While depth sensors have the potential to complement RGB data for affordance segmentation in wearable robots, their usage seems to remain underexplored.

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents

HardwareDGX agent

arXiv:2607.26076v1 Announce Type: cross Abstract: Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse c

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference

Model ReleasesDGX agent

arXiv:2607.27694v1 Announce Type: cross Abstract: Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. Ho

May have found the highest and best use case of Flux 3 - generating GPU ASMR ✨ For everyone who has been asking for access, it’s available N…

HardwareDGX agent

Justine Moore announced on July 31, 2026 that Flux 3’s latest iteration excels at generating GPU‑based ASMR content. She confirmed that this capability is now available in an early preview on the Nous

Nscale buys AI infrastructure optimization startup Anyscale for reported $1.65B

HardwareDGX agent

Data center builder Nscale Global Holdings Ltd. today announced plans to acquire Anyscale Inc., a venture-backed provider of artificial intelligence software. The terms of the deal were not disclosed.

NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek

HardwareDGX agent

NVIDIA Video Codec SDK 13.1 adds AV1 hierarchical reference mode supporting up to 31 B‑frames and efficient iterative tuning that delivers significant bitrate savings in CQ and VBR modes. It enhances

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

HardwareDGX agent

arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches prima

Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0)

Model ReleasesDGX agent

This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a

OPENAI IS GOING TO TAKE THIS ENTIRE MARKET DOWN WITH IT And you don't have to own a single share to get hurt. What I'm about to explain shou…

HardwareDGX agent

OPENAI IS GOING TO TAKE THIS ENTIRE MARKET DOWN WITH IT And you don't have to own a single share to get hurt. What I'm about to explain should worry anybody who thinks they're diversified: OpenAI is a

ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate

HardwareDGX agent

arXiv:2607.28144v1 Announce Type: cross Abstract: We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The e

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

HardwareDGX agent

arXiv:2607.28627v1 Announce Type: new Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

HardwareDGX agent

arXiv:2607.27744v1 Announce Type: new Abstract: Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far sy

Security at the foundation requires openness at the foundation. As open-weight models become critical infrastructure for the next generation…

HardwareDGX agent

Security at the foundation requires openness at the foundation. As open-weight models become critical infrastructure for the next generation of software, the systems around them must be transparent, i

Simulation of Surgical Suturing Using Position-Based Dynamics and the Material Point Method for Robot Reinforcement Learning

HardwareDGX agent

arXiv:2607.27494v1 Announce Type: new Abstract: Recent advances in robotics research have created a strong demand for high-performance simulators. Surgical robotics simulation faces unique challenges

Sources: Moonshot has a computing power agreement with Alibaba for the use of ~20K Nvidia chips; some say the deal is for H200 chips, which Alibaba denies (Mackenzie Hawkins/Bloomberg)

HardwareDGX agent

Mackenzie Hawkins / Bloomberg: Sources: Moonshot has a computing power agreement with Alibaba for the use of ~20K Nvidia chips; some say the deal is for H200 chips, which Alibaba denies — Chinese AI c

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

HardwareDGX agent

arXiv:2607.28032v1 Announce Type: new Abstract: Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian

← Previous
1…1011121314…75
Next →