AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
Model Releases

Swap models & view their capabilities! Try out in Deep Agents CLI: https://docs.langchain.com/oss/python/deepagents/cli/

DGX agent

Swap models & view their capabilities! Try out in Deep Agents CLI: https://docs.langchain.com/oss/python/deepagents/cli/ here's model profile details look like in practice, using @NVIDIAAIDev's Nemotr

model-releasesharrison-chase--x
11 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training

DGX agent

arXiv:2605.07288v1 Announce Type: cross Abstract: The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned W

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes

DGX agent

arXiv:2602.04939v2 Announce Type: replace Abstract: Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy ben

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

TAG-K: Tail-Averaged Greedy Kaczmarz for Computationally Efficient and Performant Online Inertial Parameter Estimation

DGX agent

arXiv:2510.04839v2 Announce Type: replace Abstract: Accurate online inertial parameter estimation is essential for adaptive robotic control, enabling real-time adjustment to payload changes, environme

model-releasesarxiv-cs-ro
11 May 2026
Model Releases

TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP

DGX agent

arXiv:2605.06886v1 Announce Type: new Abstract: This work introduces TajPersLexon, a curated Tajik--Persian parallel lexical resource of 40,112 word and short-phrase pairs for cross-script lexical ret

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Target-Aware Data Augmentation for SAT Prediction

DGX agent

arXiv:2605.06931v1 Announce Type: new Abstract: Learning-based approaches to NP-hard problems have shown increasing promise, but their progress is fundamentally constrained by the high cost of generat

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts

DGX agent

arXiv:2605.07256v1 Announce Type: new Abstract: Transformer architecture search (TAS) discovers optimal vision transformer (ViT) architectures automatically, reducing human effort to manually design V

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

DGX agent

arXiv:2605.07943v1 Announce Type: cross Abstract: Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple ind

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent

DGX agent

arXiv:2601.18700v2 Announce Type: replace Abstract: Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. Howeve

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Teaching Language Models to Think in Code

DGX agent

arXiv:2605.07237v1 Announce Type: new Abstract: Tool-integrated reasoning (TIR) has emerged as a dominant paradigm for mathematical problem solving in language models, combining natural language (NL)

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

TeamBench: Evaluating Agent Coordination under Enforced Role Separation

DGX agent

arXiv:2605.07073v1 Announce Type: new Abstract: Agent systems often decompose a task across multiple roles, but these roles are typically specified by prompts rather than enforced by access controls.

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Test-Time Compute Games

DGX agent

arXiv:2601.21839v2 Announce Type: replace-cross Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strate

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Text-to-CAD Evaluation with CADTests

DGX agent

arXiv:2605.07807v1 Announce Type: cross Abstract: Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The best way to level up from 1 agent => many agents. No more cycling between terminal tabs

DGX agent

This post discusses strategies for scaling from managing a single AI agent to coordinating multiple agents efficiently, likely addressing workflow challenges and tooling improvements that eliminate th

model-releasesboris-cherny--x
11 May 2026
Model Releases

The Convergence Gap: Instruction-Tuned Language Models Stabilize Later in the Forward Pass

DGX agent

arXiv:2605.07282v1 Announce Type: new Abstract: Final outputs hide when a checkpoint commits to its next-token prediction. We introduce the convergence gap, a model-diffing diagnostic that decodes eac

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

The Coupling Tax: How Shared Token Budgets Undermine Visible Chain-of-Thought Under Fixed Output Limits

DGX agent

arXiv:2605.07686v1 Announce Type: new Abstract: Chain-of-thought reasoning is often treated as a monotone way to improve language-model accuracy by letting a model think longer. We identify a counterv

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

The EDelta-MHC-Geo Transformer: Adaptive Geodesic Operations with Guaranteed Orthogonality

DGX agent

arXiv:2605.06729v1 Announce Type: cross Abstract: We present the EDelta-MHC-Geo Transformer, a novel architecture that unifies Manifold-Constrained Hyper-Connections (mHC), Deep Delta Learning (DDL),

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

DGX agent

arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

DGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking

DGX agent

arXiv:2605.06707v1 Announce Type: cross Abstract: This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the 'HTML AI B

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval

DGX agent

arXiv:2605.07186v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) benchmarks primarily focus on syntactically correct inputs, leaving a significant gap in evaluation on imperfect t

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks

DGX agent

arXiv:2605.07093v1 Announce Type: cross Abstract: The Translation Tax is often treated as a scalar: translated benchmarks are assumed to inflate scores by preserving English-source cues. We audit this

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model

DGX agent

arXiv:2602.04774v2 Announce Type: replace-cross Abstract: Setting the learning rate (LR) for a deep learning model is a critical part of successful training. Choosing LRs is often done empirically wit

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

DGX agent

arXiv:2601.23143v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

DGX agent

arXiv:2510.01290v2 Announce Type: replace Abstract: The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key-value (

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Today we’re launching the OpenAI Deployment Company to help businesses build and deploy AI. It's majority-owned and controlled by OpenAI. It…

DGX agent

Today we’re launching the OpenAI Deployment Company to help businesses build and deploy AI. It's majority-owned and controlled by OpenAI. It brings together 19 leading investment firms, consultancies,

model-releasesopenai--x
11 May 2026
Model Releases

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models

DGX agent

arXiv:2605.06683v1 Announce Type: cross Abstract: Transformer-based large language models are in some respects limited by the quadratic time and space computational complexity of attention. We introdu

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Tool Calling is Linearly Readable and Steerable in Language Models

DGX agent

arXiv:2605.07990v1 Announce Type: cross Abstract: When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. Probing 12 ins

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Tools as Continuous Flow for Evolving Agentic Reasoning

DGX agent

arXiv:2605.07339v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a s

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings

DGX agent

arXiv:2605.07158v1 Announce Type: cross Abstract: Vector search and retrieval-augmented generation (RAG) rest on the assumption that cosine similarity between text embeddings reflects conceptual relat

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion

DGX agent

arXiv:2605.07013v1 Announce Type: new Abstract: Diffusion language models (DLMs) promise parallel, order-agnostic generation, but on standard benchmarks they have historically lagged behind autoregres

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

DGX agent

arXiv:2605.07593v1 Announce Type: new Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams,

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

Tracing Uncertainty in Language Model 'Reasoning'

DGX agent

arXiv:2605.07776v1 Announce Type: cross Abstract: Language model (LM) 'reasoning', commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics u

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Traffic Scenario Orchestration from Language via Constraint Satisfaction

DGX agent

arXiv:2605.06966v1 Announce Type: new Abstract: Autonomous vehicles (AVs) require extensive testing in simulation, but test case generation for driving scenarios is laborious. The desired scenarios ar

model-releasesarxiv-cs-ro
11 May 2026
Model Releases

Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers

DGX agent

arXiv:2605.07772v1 Announce Type: new Abstract: Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically,

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

DGX agent

arXiv:2605.07924v1 Announce Type: cross Abstract: Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Dis

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

DGX agent

arXiv:2601.18744v2 Announce Type: replace Abstract: Time series are ubiquitous in real-world scenarios and crucial for applications ranging from energy management to traffic control. Consequently, the

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

UNCOM: Zero-shot Context-Aware Command Understanding for Tabletop Scenarios

DGX agent

arXiv:2410.06355v3 Announce Type: replace-cross Abstract: This paper presents UNCOM, a novel hybrid framework for interpreting natural human commands in tabletop scenarios. The system integrates multi

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Understanding Robustness of Model Editing in Code LLMs

DGX agent

arXiv:2511.03182v2 Announce Type: replace-cross Abstract: Large language models (LLMs) for code are increasingly used in software development, but they remain static after pretraining while APIs and s

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Uneven Evolution of Cognition Across Generations of Generative AI Models

DGX agent

arXiv:2605.06815v1 Announce Type: new Abstract: The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task perfor

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts

DGX agent

arXiv:2605.07395v1 Announce Type: cross Abstract: Efficient routing across multiple LLMs enables cost-quality tradeoffs by directing queries to the cheapest capable model. Prior work attributes routin

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Using LLM in the shebang line of a script

DGX agent

TIL: Using LLM in the shebang line of a script Kim_Bruning on Hacker News: But seriously, you can put a shebang on an english text file now (if you're sufficiently brave) [...] This inspired me to loo

model-releasessimon-willison
11 May 2026
Model Releases

Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset

DGX agent

arXiv:2602.16571v2 Announce Type: replace Abstract: Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrie

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

v0.30.0-rc14: Merge remote-tracking branch 'upstream/main' into llama-runner-phase-0

DGX agent

This is a release candidate version of Ollama that merges updates from the main development branch into the llama-runner-phase-0 branch, likely incorporating bug fixes and features in preparation for

model-releasesollama-releases
11 May 2026
Model Releases

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

DGX agent

arXiv:2605.07872v1 Announce Type: cross Abstract: Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely l

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Visual Text Compression as Measure Transport

DGX agent

arXiv:2605.06708v1 Announce Type: cross Abstract: Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language mod

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts

DGX agent

arXiv:2605.06175v2 Announce Type: replace Abstract: Vision-language-action (VLA) models inherit rich visual-semantic priors from pre-trained vision-language backbones, but adapting them to robotic con

model-releasesarxiv-cs-ro
11 May 2026
Model Releases

📣We're calling for ambassadors! Whether you're a developer with great technical taste or a local community leader who loves bringing people…

DGX agent

📣We're calling for ambassadors! Whether you're a developer with great technical taste or a local community leader who loves bringing people together, we'd love to have you join us. Visit the website b

model-releasesqwen--x
11 May 2026
← Previous
1…343344345346347…471
Next →