AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,732 results
6 May 2026

ZeRO-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

HardwareDGX agent

arXiv:2605.02960v1 Announce Type: new Abstract: Production LLM workloads increasingly serve discriminative tasks, such as classification, recommendation, and verification, whose answers are read from

5 May 2026

A Semantic Autonomy Framework for VLM-Integrated Indoor Mobile Robots: Hybrid Deterministic Reasoning and Cross-Robot Adaptive Memory

HardwareDGX agent

arXiv:2605.02525v1 Announce Type: new Abstract: Autonomous indoor mobile robots can navigate reliably to metric coordinates using established frameworks such as ROS 2 Navigation 2, yet they lack the a

aerial-autonomy-stack -- a Faster-than-real-time, Autopilot-agnostic, ROS2 Framework to Simulate and Deploy Perception-based Drones


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
HardwareDGX agent

arXiv:2602.07264v2 Announce Type: replace Abstract: Unmanned aerial vehicles are rapidly transforming multiple applications, from agricultural and infrastructure monitoring to logistics and defense. I

AgriKD: Cross-Architecture Knowledge Distillation for Efficient Leaf Disease Classification

HardwareDGX agent

arXiv:2605.01355v1 Announce Type: new Abstract: Automated leaf disease classification is critical for early disease detection in resource-constrained field environments. Vision Transformers (ViTs) pro

Blitzy raises 200M at 1.4B valuation to deploy thousands of coding agents in parallel

HardwareDGX agent

Autonomous software development startup Blitzy Inc. said today it has raised 200 million in new funding on a valuation of 1.4 billion to expand its enterprise coding platform. The company was founded

Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness

HardwareDGX agent

arXiv:2605.01006v1 Announce Type: new Abstract: Partisan news media erode cross-partisan trust, but large language models (LLMs) offer a potential means of debiasing such content at scale. Across two

Deepinfra lands $107M in funding to build out its dedicated inference cloud for open-source models

HardwareDGX agent

Dedicated inference cloud startup Deepinfra Inc. is looking to expand its global capacity after raising 107 million in a Series B round of funding led by 500 Global and Georges Harik, who was one of G

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning

HardwareDGX agent

arXiv:2510.09883v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve state-of-the-art performance on challenging benchmarks by generating long chains of intermediate steps, but th

Effective Capacitance Modeling Using Graph Neural Networks

HardwareDGX agent

arXiv:2507.03787v2 Announce Type: replace Abstract: Static timing analysis is a crucial stage in the VLSI design flow that verifies the timing correctness of circuits. Timing analysis depends on the p

Fast Log-Domain Sinkhorn Optimal Transport with Warp-Level GPU Reductions

HardwareDGX agent

arXiv:2605.00837v1 Announce Type: new Abstract: Entropic regularized optimal transport (OT) via the Sinkhorn algorithm has become a fundamental tool in machine learning, yet existing implementations e

How to Build In-Vehicle AI Agents with NVIDIA: From Cloud to Car

HardwareDGX agent

The automotive cockpit is undergoing a fundamental shift from rule-based interfaces to agentic, multimodal AI systems capable of reasoning, planning, and acting. This is enabled by an agentic AI pipel

Inference is 80-90% of the lifetime cost of a production AI system. Most AI-native teams are leaving performance and margin on the table. He…

HardwareDGX agent

Inference is 80-90% of the lifetime cost of a production AI system. Most AI-native teams are leaving performance and margin on the table. Here’s how Together AI, the AI Native Cloud, fixes that on @nv

InfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain Generation

HardwareDGX agent

arXiv:2512.08309v4 Announce Type: replace Abstract: For decades, procedural worlds have been built on procedural noise functions such as Perlin noise, which are fast and infinite, yet fundamentally li

Lambda assembles leadership team to power gigawatt-scale AI infrastructure for the superintelligence era

HardwareDGX agent

Lambda Labs has assembled a new leadership team to develop and manage large-scale AI infrastructure capable of supporting gigawatt-level power consumption for advanced AI models. The company is positi

LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference

HardwareDGX agent

arXiv:2605.01058v1 Announce Type: cross Abstract: Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; ye

LLM Output Detectability and Task Performance Can be Jointly Optimized

HardwareDGX agent

arXiv:2605.01350v1 Announce Type: new Abstract: Detecting machine-generated text is essential for transparency and accountability when deploying large language models (LLMs). Among detection approache

Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models

HardwareDGX agent

arXiv:2605.01870v1 Announce Type: new Abstract: Large Language Models (LLMs) have substantially advanced the field of Natural Language Processing (NLP), achieving state-of-the-art performance across a

MusicInfuser: Making Video Diffusion Listen and Dance

HardwareDGX agent

arXiv:2503.14505v3 Announce Type: replace Abstract: We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized wit

Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion

HardwareDGX agent

arXiv:2603.10584v2 Announce Type: replace Abstract: We introduce Marigold-SSD, a single-step, late-fusion depth completion framework that leverages strong diffusion priors while eliminating the costly

No GPU utilization

HardwareDGX agent

Ollama not using GPU is commonly diagnosed by running `ollama ps`—if it shows 100% CPU, Ollama isn't detecting the GPU. The issue typically has multiple causes, with common ones including driver probl

Nscale agrees to spend €695M on infrastructure in Portugal in an expansion of its Microsoft partnership, supplying 66K+ Nvidia Rubin GPUs from late 2027 (Paula Doenecke/Bloomberg)

HardwareDGX agent

Paula Doenecke / Bloomberg: Nscale agrees to spend €695M on infrastructure in Portugal in an expansion of its Microsoft partnership, supplying 66K+ Nvidia Rubin GPUs from late 2027 — Nscale Global Hol

Nvidia, AMD back $100M round for AI tooling startup RadixArk

HardwareDGX agent

RadixArk Inc., a startup that provides tools for artificial intelligence developers, has raised 100 million from a group of high-profile backers. Nvidia Corp.’s NVentures fund led the seed round with

NVIDIA and ServiceNow Partner on New Autonomous AI Agents for Enterprises

HardwareDGX agent

Enterprise AI has learned to generate. It has learned to reason. Now companies are asking the next question: How should AI act? Early agent systems have shown what’s possible, moving beyond simple pro

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization

HardwareDGX agent

arXiv:2602.02958v4 Announce Type: replace Abstract: Despite rapid progress in autoregressive video diffusion, an emerging system algorithm bottleneck limits both deployability and generation capabilit

Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

HardwareDGX agent

arXiv:2605.00254v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of

Silicon Valley bets $200M on AI data centers floating in the ocean

HardwareDGX agent

Panthalassa, a US startup, is developing floating data centers powered by ocean waves and cooled by seawater to address AI infrastructure's growing energy demands. Peter Thiel led a $140 million inves

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving

HardwareDGX agent

arXiv:2605.01708v1 Announce Type: cross Abstract: Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to better load-balance between the compute-bound

Stochastic Sparse Attention for Memory-Bound Inference

HardwareDGX agent

arXiv:2605.01910v1 Announce Type: new Abstract: Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all n_k key and value vectors from KV cach

The extit{Silicon Society} Cookbook: Design Space of LLM-based Social Simulations

HardwareDGX agent

arXiv:2605.00197v1 Announce Type: cross Abstract: Studies attempting to simulate human behavior with extit{Silicon Societies} grow in numbers while LLM-only social networks have started appearing outs

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning

HardwareDGX agent

arXiv:2602.13595v2 Announce Type: replace Abstract: Neural scaling laws provide a predictable recipe for AI advancement: reducing numerical precision should linearly improve computational efficiency a

ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA

HardwareDGX agent

arXiv:2605.01935v1 Announce Type: cross Abstract: Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs),

We cut model cold-start times from minutes to seconds. 60x faster, saving thousands of GPU-minutes and hundreds of terabytes of transfers ev…

HardwareDGX agent

We cut model cold-start times from minutes to seconds. 60x faster, saving thousands of GPU-minutes and hundreds of terabytes of transfers every day. It turns out GPUs already holding weights are faste

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

HardwareDGX agent

arXiv:2605.02262v1 Announce Type: cross Abstract: Recently, video language models (VLMs) have been applied in various fields. However, the visual token sequence of the VLM is too long, which may cause

4 May 2026

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments

HardwareDGX agent

arXiv:2605.00650v1 Announce Type: new Abstract: Fine-tuning LLMs is necessary for various dedicated downstream tasks, but classic backpropagation-based fine-tuning methods require substantial GPU memo

Boosting multimodal inference performance by >10% with a single Python dictionary

HardwareDGX agent

This Modal blog post describes a performance optimization technique for multimodal AI inference that achieves over 10% improvement through a simple Python dictionary-based approach. The article likely

Cerebras seeks a valuation of up to 26.62B in its US IPO, aiming to raise 3.5B by selling 28M shares at 115 to 125 apiece in its second attempt to go public (Reuters)

HardwareDGX agent

Reuters: Cerebras seeks a valuation of up to 26.62B in its US IPO, aiming to raise 3.5B by selling 28M shares at 115 to 125 apiece in its second attempt to go public — Nvidia-rival Cerebras is seeking

Down payment on a EUV machine!

HardwareDGX agent

Dylan Patel likely announced or commented on a significant financial commitment or investment milestone related to an EUV (Extreme Ultraviolet) lithography machine, which are critical semiconductor ma

DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference

HardwareDGX agent

arXiv:2605.00174v1 Announce Type: cross Abstract: Video and image streaming on edge devices requires low latency. To address this, Neural Networks (NNs) are widely used, and prior work mainly focuses

Lattice Semiconductor agrees to acquire AMI for $1.65B in cash and stock; Georgia-based AMI provides firmware and infrastructure manageability for cloud and AI (Mike Rogoway/Oregonian)

HardwareDGX agent

Mike Rogoway / Oregonian: Lattice Semiconductor agrees to acquire AMI for 1.65B in cash and stock; Georgia-based AMI provides firmware and infrastructure manageability for cloud and AI — Hillsboro-bas

Legislators and experts criticize the EU's €20B sovereign compute data center plan, questioning whether there is demand and the plan's reliance on Nvidia GPUs (Pieter Haeck/Politico)

HardwareDGX agent

Pieter Haeck / Politico: Legislators and experts criticize the EU's €20B sovereign compute data center plan, questioning whether there is demand and the plan's reliance on Nvidia GPUs — Critics argue

llamacpp on Apple Silicon, once configured correctly is really rock solid! You can throw anything at it and it will answer. Impressive!

HardwareDGX agent

Llamacpp, when properly configured on Apple Silicon hardware, demonstrates robust performance and reliability for running language models. The tool can handle varied input requests effectively, making

MiniVLA-Nav v1: A Multi-Scene Simulation Dataset for Language-Conditioned Robot Navigation

HardwareDGX agent

arXiv:2605.00397v1 Announce Type: new Abstract: We present MiniVLA-Nav v1, a simulation dataset for Language-Conditioned Object Approach (LCOA) navigation: given a short natural-language instruction,

Most AI teams treat compute as a commodity. It's not.

HardwareDGX agent

AI compute resources should not be treated as interchangeable commodities, as different hardware configurations, providers, and architectures significantly impact training costs, inference latency, an

Optimize Supply Chain Decision Systems Using NVIDIA cuOpt Agent Skills

HardwareDGX agent

NVIDIA cuOpt Agent Skills integrate LLM reasoning with GPU-accelerated solvers to enable AI agents to translate natural language supply chain problems into optimized mathematical models, with skills e

The case against OpenAI is getting markedly stronger now that Musk is off the stand. Why? Musk’s lawyer is interrogating OpenAI founder Greg…

HardwareDGX agent

The case against OpenAI is getting markedly stronger now that Musk is off the stand. Why? Musk’s lawyer is interrogating OpenAI founder Greg Brockman, making clear that OpenAI sold its mission as a no

3 May 2026

Analysis: Asian suppliers account for ~90% of Nvidia's production costs, up from 65% in 2025, as latest wave of collaborations shifts from chips to physical AI (Abhishek Vishnoi/Bloomberg)

HardwareDGX agent

Abhishek Vishnoi / Bloomberg: Analysis: Asian suppliers account for ~90% of Nvidia's production costs, up from 65% in 2025, as latest wave of collaborations shifts from chips to physical AI — The list

Are babies and relationships measured on the same time scale? After 18-24 months you start measuring in years.

HardwareDGX agent

This post draws a comparison between how parents measure their babies' development and relationship milestones, noting that both shift from month-based to year-based metrics around the 18-24 month mar

NVIDIA's New AI Turns One Photo Into A World That Never Breaks

HardwareDGX agent

NVIDIA's Instant Splat is a breakthrough AI technique that synthesizes missing information from a few images to create immersive 3D models of virtual worlds. The method produces high-quality 3D models

People hate on SF men for group think, but every girl in SF has Tabi's now It's an epidemic

HardwareDGX agent

This post humorously critiques conformity in San Francisco by observing that while men in the tech scene are often stereotyped as groupthink followers, women in SF have similarly adopted Tabis (Japane

The number of jobs in the future is endless because the problems to solve are endless. Jobs multiply as we get more complex. No AI or human …

HardwareDGX agent

The number of jobs in the future is endless because the problems to solve are endless. Jobs multiply as we get more complex. No AI or human can solve all problems and all the work to do in the Univers

torch-nvenc-compress: GPU NVENC silicon as a PCIe bandwidth multiplier — PCA + pure-ctypes Video Codec SDK wrapper. Parallel-path overlap measured at 67% of theoretical max on a real GEMM + encode workload. [P]

HardwareDGX agent

torch-nvenc-compress is a Python library that leverages GPU NVENC (NVIDIA's hardware video encoding) to optimize PCIe bandwidth utilization by compressing data during transfer. The project implements

Whenever I make a group chat green, I now get to proudly proclaim it as goblin mode

HardwareDGX agent

The post humorously references Apple's iMessage feature that displays group chats in green when they include non-iPhone users, joking that activating 'goblin mode' (a playful term for behaving mischie

2 May 2026

CUDA V.13?

HardwareDGX agent

Ollama's MLX engine runs on NVIDIA GPUs via CUDA v13 on Windows and Linux. Users have reported that Ollama crashes on RTX 3060 with cuda_v13 in versions 0.13.0 through 0.15.6, while deleting the cuda_

Intel Inside the Micro Revolution: 8008 Origins

HardwareDGX agent

This article explores the origins and development of Intel's 8008 microprocessor, a pioneering chip that played a crucial role in the early microcomputer revolution of the 1970s. It likely examines th

Sources: Anthropic is in early talks to buy AI inference chips from UK-based Fractile when they become available in 2027 (The Information)

HardwareDGX agent

The Information: Sources: Anthropic is in early talks to buy AI inference chips from UK-based Fractile when they become available in 2027 — As Anthropic's sales explode, straining the servers it uses,

1 May 2026

AI Value Capture - The Shift To Model Labs

HardwareDGX agent

This article examines how value in the AI industry is shifting away from traditional chip manufacturers toward model development labs and companies that build large language models. It likely analyzes

AI Value Capture - The Shift To Model Labs Vera Rubin VR NVL72: V for Value - Rubin delivers a step jump in performance per TCO. ROI accruin…

HardwareDGX agent

AI Value Capture - The Shift To Model Labs Vera Rubin VR NVL72: V for Value - Rubin delivers a step jump in performance per TCO. ROI accruing to users, Neoclouds, Hyperscalers, AI Labs, Memory Vendors

All the Doomers and hawks are lining up behind this distillation 'attack' farce because they want to see open source banned. It's really as …

HardwareDGX agent

All the Doomers and hawks are lining up behind this distillation 'attack' farce because they want to see open source banned. It's really as simple as that. They want to take away your right to choose,

Benchmarking Deep Learning Models for Object Detection on Edge Computing Devices

HardwareDGX agent

arXiv:2409.16808v1 Announce Type: cross Abstract: Modern applications, such as autonomous vehicles, require deploying deep learning algorithms on resource-constrained edge devices for real-time image

By popular request, YOLO-mode has landed on ml-intern https://smolagents-ml-intern.hf.space This enables the intern to execute long-running …

HardwareDGX agent

By popular request, YOLO-mode has landed on ml-intern https://smolagents-ml-intern.hf.space This enables the intern to execute long-running tasks like launching parallel ablations to determine the opt

← Previous
1…2122232425…29
Next →