AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,732 results
4 Jun 2026

Cartridges at Scale: Training Modular KV Caches over Large Document Collections

HardwareDGX agent

arXiv:2606.04557v1 Announce Type: new Abstract: Large Language Models can reason over long contexts, yet prefilling millions of tokens is wasteful as much of the content remains static across queries.

DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation

HardwareDGX agent

arXiv:2606.04432v1 Announce Type: new Abstract: Video diffusion transformers have achieved state-of-the-art visual quality, but their high inference cost remains a major bottleneck for real-time appli

Ex-OpenAI Tech Lead, Justin Lebar joins SemiAnalysis as an Visiting Fellow to Burn $10,000 in 3 hours to find dozens of AMDGPU LLVM, x86 LLV…

HardwareDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Ex-OpenAI Tech Lead, Justin Lebar joins SemiAnalysis as an Visiting Fellow to Burn $10,000 in 3 hours to find dozens of AMDGPU LLVM, x86 LLVM, NVPTX bugs 00:00 - Intro & Justin’s background 00:59 - Ho

Fitting scattered data with optional monotonicity constraints on GPU: LipFit package

HardwareDGX agent

arXiv:2606.04670v1 Announce Type: cross Abstract: This paper presents a method of multivariate scattered data interpolation and approximation that produces optimal Lipschitz-continuous approximation,

Forecast: Fun Ahead — 18 Games Join in June to Stream on GeForce NOW

HardwareDGX agent

June’s forecast with GeForce NOW: 100% chance of gaming. GeForce NOW is lining up new adventures for the month, from big-name blockbusters to quirky indies ready for the spotlight. Members can dive in

Generalist AI raises 400M at 2B valuation to build general intelligence for robotics

HardwareDGX agent

Artificial intelligence startup Generalist AI Inc., a startup building embodied robotics intelligence, said today it has raised 400 million in new funding, bringing the company’s valuation to 2 billio

I love that people reposting what we said miss most of what the note says. Happens all the time.

HardwareDGX agent

Dylan Patel observes that when people repost or share his notes on social media, they frequently misrepresent or overlook the main points of his original message. He notes this is a recurring pattern

LLM Compression with Jointly Optimizing Architectural and Quantization choices

HardwareDGX agent

arXiv:2606.04063v1 Announce Type: cross Abstract: Deploying large language models (LLMs) is challenging due to their significant memory and computational requirements. While some methods address this

Nvidia snaps up Kumo AI, a predictive AI startup known for its extreme accuracy

HardwareDGX agent

Nvidia Corp. has bagged itself another artificial intelligence startup, acquiring four-year-old model maker Kumo AI Inc. The company designs AI models focused on making extremely accurate business pre

OpenAI is in deep, deep trouble They are low on capital relative to the massive cash they are burning and there is only so much capital in t…

HardwareDGX agent

OpenAI is in deep, deep trouble They are low on capital relative to the massive cash they are burning and there is only so much capital in the world. No rational person would sell their Bitcoin or Nvi

Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models

HardwareDGX agent

arXiv:2606.04287v1 Announce Type: cross Abstract: Generating realistic and diverse graphs is a key problem in machine learning, with applications in molecular discovery, circuit design, cybersecurity,

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

HardwareDGX agent

arXiv:2606.04511v1 Announce Type: new Abstract: Sparse attention reduces compute and memory bandwidth for long-context LLM inference. However, two key challenges remain: (1) KV cache capacity still gr

TITAN-FedAnil+: Trust-Based Adaptive Blockchain Federated Learning for Resource-Constrained Intelligent Enterprises

HardwareDGX agent

arXiv:2606.04388v1 Announce Type: cross Abstract: Federated Learning (FL) has emerged as an effective paradigm for collaborative intelligence while preserving data privacy. However, data heterogeneity

Together AI provides the inference stack behind both: high-throughput serving on the latest NVIDIA Blackwell GPUs for agentic workloads, and…

HardwareDGX agent

Together AI provides the inference stack behind both: high-throughput serving on the latest NVIDIA Blackwell GPUs for agentic workloads, and TensorRT engines plus event-driven streaming I/O for low-la

3 Jun 2026

Announcing Foundry Managed Compute: Run open models in Microsoft Foundry

HardwareDGX agent

Microsoft Foundry Managed Compute is a new GPU platform-as-a-service for hosting open-source and custom AI models behind the same endpoint, SDKs, and bill as frontier models. The post Announcing Found

CoreWeave’s Vera Rubin milestone sets stage for theCUBE’s agentic AI coverage

HardwareDGX agent

The agentic AI era is putting new pressure on the infrastructure stack, and CoreWeave Inc.’s latest milestone gives the conversation a sharper edge. This week, the company announced that it has comple

CRAM-ER: Error-Resilient Spintronic Computational Random Access Memory for Scalable In-Memory Computation

HardwareDGX agent

arXiv:2606.02781v1 Announce Type: cross Abstract: Deep neural networks (DNNs) have achieved state-of-the-art performance across diverse domains. However, typical Von Neumann compute paradigms face sev

Floating Point: The Origin Story

HardwareDGX agent

This article explores the historical development and origins of floating-point number representation in computing, likely covering how early computer scientists and engineers designed methods to repre

GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning

HardwareDGX agent

arXiv:2606.02857v1 Announce Type: cross Abstract: Zeroth-order (ZO) optimization is a memory-efficient alternative to backpropagation for fine-tuning large language models, but its deployment is limit

MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference Concurrency

HardwareDGX agent

arXiv:2606.03014v1 Announce Type: new Abstract: Mixture-of-Agents (MoA) systems improve reasoning accuracy by routing each query to multiple expert LLMs and aggregating their outputs. Efficiently exec

Nvidia acquired Kumo, which sells predictive AI software to enterprises, a source says for 400M+; PitchBook: Kumo raised 37M at a $250M valuation in 2022 (The Information)

HardwareDGX agent

The Information: Nvidia acquired Kumo, which sells predictive AI software to enterprises, a source says for 400M+; PitchBook: Kumo raised 37M at a 250M valuation in 2022 — Nvidia has bought Kumo AI, a

NVIDIA Enables the Next Era Of Physical AI Research With Agent Skills For Autonomous Vehicles, Robotics And Vision AI

HardwareDGX agent

At CVPR, NVIDIA is unveiling new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots and vision AI systems. The core challenge in physica

NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics

HardwareDGX agent

arXiv:2606.03551v1 Announce Type: new Abstract: Simulation has become a core infrastructure for robotics research. Unlike previous simulators, NVIDIA Isaac Sim leverages GPU acceleration to enable lar

NVIDIA Research Unlocks Advanced Grasping, Smarter Autonomous Driving and Agent Training at Scale

HardwareDGX agent

What makes a robot gripper useful isn’t that it can pick up one object — it’s that it can pick up the next one, and the one after that, with a tool it’s never held before. What makes an autonomous veh

Pruning Deep Neural Networks via the Marchenko--Pastur Distribution

HardwareDGX agent

arXiv:2606.02608v1 Announce Type: new Abstract: We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets. The main prac

Quobly raises $150M for its silicon spin qubit technology

HardwareDGX agent

French quantum computing startup Quobly SAS today announced that it has raised €130 million, or about 150 million, to commercialize its technology. The Series A round was led by publicly traded chipma

Releasing vui an open source voice mode 300M TTS model Runs on a single consumer gpu / apple sillicon Context aware speech 6 minutes of cont…

HardwareDGX agent

Jeremy Howard announced the release of Vui, an open-source voice mode text-to-speech (TTS) model with 300 million parameters that can run on consumer GPUs and Apple Silicon. The model features context

Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles

HardwareDGX agent

arXiv:2505.08222v3 Announce Type: replace-cross Abstract: Autonomous vehicles (AVs) offer a cost-effective solution for scientific missions such as underwater tracking. Reinforcement learning (RL) has

Speedrunning Tabular Foundation Model Pretraining

HardwareDGX agent

arXiv:2606.03681v1 Announce Type: new Abstract: Pretraining cost is a major bottleneck for research on tabular foundation models, slowing the iteration cycle for new architectures, priors, and optimiz

To Boldly Go: The Case for Space Datacenters

HardwareDGX agent

This article explores the potential viability and benefits of locating data centers in space rather than on Earth, likely discussing advantages such as reduced cooling costs due to the vacuum environm

Towards Compact Autonomous Driving Perception with Balanced Learning and Multi-sensor Fusion

HardwareDGX agent

arXiv:2606.02979v1 Announce Type: cross Abstract: We present a novel compact deep multi-task learning model to handle various autonomous driving perception tasks in one forward pass. The model perform

What’s new in serverless Managed Service for Apache Spark

HardwareDGX agent

Whether you use it for data preparation, real-time interactive queries, AI model training, or something entirely different, running Apache Spark at scale is demanding — you shouldn’t have to manage th

2 Jun 2026

APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention

HardwareDGX agent

arXiv:2601.21444v2 Announce Type: replace-cross Abstract: The efficiency of long-video inference remains a critical bottleneck, mainly due to the dense computation in the prefill stage of Large Multim

Bit-Exact AI Inference Verification Without Performance Tradeoffs

HardwareDGX agent

arXiv:2606.00279v1 Announce Type: cross Abstract: Verifying claims about AI workloads is a pre- requisite for credible AI governance of covert adversaries (who comply with monitoring only when detecti

BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding

HardwareDGX agent

arXiv:2606.00144v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel. In resourc

Cohort-Scale Neural Atlases of Ultrasound Video

HardwareDGX agent

arXiv:2606.00890v1 Announce Type: new Abstract: Ultrasound is the most widely used real-time imaging modality in clinical practice, yet per-frame video annotation remains a major bottleneck: expert la

CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving

HardwareDGX agent

arXiv:2603.28768v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) has recently emerged as the mainstream architecture for efficiently scaling large language models while maintaining n

Deploy Agentic-Ready AI at the Edge with Memory Efficiency in NVIDIA JetPack 7.2

HardwareDGX agent

NVIDIA JetPack 7.2 enables deployment of AI agents to edge devices with optimized memory and performance for real-world applications. The release directly supports one-command deployment of NVIDIA Nem

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap

HardwareDGX agent

arXiv:2512.10236v2 Announce Type: replace-cross Abstract: Modern ML workloads demand distributing training and inference across multiple GPUs. However, these parallelization techniques often suffer fr

Don't Let a Few Network Failures Slow the Entire AllReduce

HardwareDGX agent

arXiv:2606.01680v1 Announce Type: cross Abstract: Network failures are among the most frequent hardware faults in large-scale GPU clusters and a leading cause of training-job interruptions. Modern col

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

HardwareDGX agent

arXiv:2511.04791v2 Announce Type: replace Abstract: Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive pr

Eyettention II: A Dual-Sequence Architecture for Modeling Fixation Location, Within-Word Landing Position, and Fixation Duration in Reading

HardwareDGX agent

arXiv:2606.01964v1 Announce Type: new Abstract: The way our eyes move while reading provides valuable insights into both the reader's cognitive processes and the properties of the text. In particular,

First technical Deepdive on M3 on the internet😎

HardwareDGX agent

First technical Deepdive on M3 on the internet😎 MiniMax-M3 combines 1M context, native multimodality, and MiniMax Sparse Attention. The next layer is serving it efficiently: KV-block-major sparse atte

FLARE: Diffusion for Hybrid Language Model

HardwareDGX agent

arXiv:2606.01774v1 Announce Type: cross Abstract: Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck for low-laten

FreqLite: A Lightweight Frequency-Decomposed Linear Model with Adaptive Reversible Normalization for Robust Long-Term Time-Series Forecasting

HardwareDGX agent

arXiv:2606.01339v1 Announce Type: cross Abstract: Long-term time-series forecasting needs models that are accurate yet efficient enough for commodity hardware. Lightweight linear forecasters are remar

Industrial Software Leaders Build Secure, Autonomous AI Engineers With NVIDIA NemoClaw

HardwareDGX agent

Accelerated computing has revolutionized industrial engineering, compressing simulation times from weeks to hours. Today’s remaining challenges sit in the end-to-end workflow surrounding the simulatio

Introducing workspaces for Lambda Cloud

HardwareDGX agent

Lambda workspaces help teams organize cloud resources, control access, and separate dev, staging, and production in shared GPU environments. A junior researcher kills a production training run. A cont

Join the livestream to hear from our team members @karan4d and @yoniebans live from @nvidia! https://www.youtube.com/watch?v=pgQDbRMa2Eg

HardwareDGX agent

Nous Research is hosting a livestream featuring team members Karan and Yoni speaking from NVIDIA, likely discussing AI research, model development, or collaboration between Nous Research and NVIDIA. T

LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models

HardwareDGX agent

arXiv:2606.01838v1 Announce Type: cross Abstract: Agentic language model systems alternate between two structurally distinct step types: structured tool calls (short, deterministic, low perplexity) an

LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching

HardwareDGX agent

arXiv:2606.00228v1 Announce Type: new Abstract: In semiconductor manufacturing, lithography projects circuit layouts onto silicon wafers through an optical mask. As circuit features shrink below the w

Lodestar: An Online-Learning LLM Inference Router

HardwareDGX agent

arXiv:2606.00946v1 Announce Type: cross Abstract: Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for G

Marvell shares closed up 32.52% on Tuesday after Jensen Huang hailed the chipmaker as the 'next trillion-dollar company' at Computex (Sawdah Bhaimiya/CNBC)

HardwareDGX agent

Sawdah Bhaimiya / CNBC: Marvell shares closed up 32.52% on Tuesday after Jensen Huang hailed the chipmaker as the “next trillion-dollar company” at Computex — Nvidia's CEO Jensen Huang hailed Marvell

Microsoft announces Surface RTX Spark AI supercomputer development box

HardwareDGX agent

Microsoft Corp. today announced it’s bringing agentic development capabilities into the hands of developers with a new desktop form factor supercomputer called the Surface RTX Spark Dev Box. Developed

MiniMax-M3 combines 1M context, native multimodality, and MiniMax Sparse Attention. The next layer is serving it efficiently: KV-block-major…

HardwareDGX agent

MiniMax-M3 combines 1M context, native multimodality, and MiniMax Sparse Attention. The next layer is serving it efficiently: KV-block-major sparse attention, paged MSA decode, optimized index scoring

NVIDIA Jetson Brings Agentic AI to the Physical World

HardwareDGX agent

Agentic AI is getting physical. At COMPUTEX on Tuesday, NVIDIA announced NVIDIA JetPack 7.2 and NVIDIA NemoClaw support on NVIDIA Jetson. JetPack 7.2 brings agentic AI skills, Yocto project support, N

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving

HardwareDGX agent

arXiv:2606.01839v1 Announce Type: cross Abstract: LLM-based agents resolve a user task through many turns of dependent inference and tool calls, producing a workload whose total cost is unknown when t

Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era

HardwareDGX agent

NVIDIA and DigitalOcean are collaborating to develop an open infrastructure stack designed to support autonomous AI agents that operate continuously. The initiative emphasizes open-source principles a

Opus 4.8

HardwareDGX agent

Opus 4.8 likely refers to a software release, update, or feature announcement covered by Ben's Bites, a technology newsletter focused on AI and software developments. Without access to the specific ar

PortBERT: Navigating the Depths of Portuguese Language Models

HardwareDGX agent

arXiv:2606.02100v1 Announce Type: new Abstract: Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neg

Practical Aspects on Solving Differential Equations Using Deep Learning: A Primer

HardwareDGX agent

arXiv:2408.11266v5 Announce Type: replace Abstract: Deep learning is now common across many scientific fields, including the study of partial differential equations. This article provides a brief, acc

← Previous
1…1213141516…29
Next →