AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,555 results
11 May 2026

Structured Prototype-Guided Adaptation for EEG Foundation Models

Model ReleasesDGX agent

arXiv:2602.17251v2 Announce Type: replace Abstract: Electroencephalography (EEG) foundation models (EFMs) have shown strong potential for transferable representation learning, yet their adaptation in

Swap models & view their capabilities! Try out in Deep Agents CLI: https://docs.langchain.com/oss/python/deepagents/cli/

Model ReleasesDGX agent

Swap models & view their capabilities! Try out in Deep Agents CLI: https://docs.langchain.com/oss/python/deepagents/cli/ here's model profile details look like in practice, using @NVIDIAAIDev's Nemotr

Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.07288v1 Announce Type: cross Abstract: The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned W

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes

Model ReleasesDGX agent

arXiv:2602.04939v2 Announce Type: replace Abstract: Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy ben

TAG-K: Tail-Averaged Greedy Kaczmarz for Computationally Efficient and Performant Online Inertial Parameter Estimation

Model ReleasesDGX agent

arXiv:2510.04839v2 Announce Type: replace Abstract: Accurate online inertial parameter estimation is essential for adaptive robotic control, enabling real-time adjustment to payload changes, environme

TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP

Model ReleasesDGX agent

arXiv:2605.06886v1 Announce Type: new Abstract: This work introduces TajPersLexon, a curated Tajik--Persian parallel lexical resource of 40,112 word and short-phrase pairs for cross-script lexical ret

Target-Aware Data Augmentation for SAT Prediction

Model ReleasesDGX agent

arXiv:2605.06931v1 Announce Type: new Abstract: Learning-based approaches to NP-hard problems have shown increasing promise, but their progress is fundamentally constrained by the high cost of generat

TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts

Model ReleasesDGX agent

arXiv:2605.07256v1 Announce Type: new Abstract: Transformer architecture search (TAS) discovers optimal vision transformer (ViT) architectures automatically, reducing human effort to manually design V

TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

Model ReleasesDGX agent

arXiv:2605.07943v1 Announce Type: cross Abstract: Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple ind

TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent

Model ReleasesDGX agent

arXiv:2601.18700v2 Announce Type: replace Abstract: Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. Howeve

Teaching Language Models to Think in Code

Model ReleasesDGX agent

arXiv:2605.07237v1 Announce Type: new Abstract: Tool-integrated reasoning (TIR) has emerged as a dominant paradigm for mathematical problem solving in language models, combining natural language (NL)

TeamBench: Evaluating Agent Coordination under Enforced Role Separation

Model ReleasesDGX agent

arXiv:2605.07073v1 Announce Type: new Abstract: Agent systems often decompose a task across multiple roles, but these roles are typically specified by prompts rather than enforced by access controls.

Test-Time Compute Games

Model ReleasesDGX agent

arXiv:2601.21839v2 Announce Type: replace-cross Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strate

Text-to-CAD Evaluation with CADTests

Model ReleasesDGX agent

arXiv:2605.07807v1 Announce Type: cross Abstract: Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there

The best way to level up from 1 agent => many agents. No more cycling between terminal tabs

Model ReleasesDGX agent

This post discusses strategies for scaling from managing a single AI agent to coordinating multiple agents efficiently, likely addressing workflow challenges and tooling improvements that eliminate th

The Convergence Gap: Instruction-Tuned Language Models Stabilize Later in the Forward Pass

Model ReleasesDGX agent

arXiv:2605.07282v1 Announce Type: new Abstract: Final outputs hide when a checkpoint commits to its next-token prediction. We introduce the convergence gap, a model-diffing diagnostic that decodes eac

The Coupling Tax: How Shared Token Budgets Undermine Visible Chain-of-Thought Under Fixed Output Limits

Model ReleasesDGX agent

arXiv:2605.07686v1 Announce Type: new Abstract: Chain-of-thought reasoning is often treated as a monotone way to improve language-model accuracy by letting a model think longer. We identify a counterv

The EDelta-MHC-Geo Transformer: Adaptive Geodesic Operations with Guaranteed Orthogonality

Model ReleasesDGX agent

arXiv:2605.06729v1 Announce Type: cross Abstract: We present the EDelta-MHC-Geo Transformer, a novel architecture that unifies Manifold-Constrained Hyper-Connections (mHC), Deep Delta Learning (DDL),

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

Model ReleasesDGX agent

arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

Model ReleasesDGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking

Model ReleasesDGX agent

arXiv:2605.06707v1 Announce Type: cross Abstract: This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the 'HTML AI B

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval

Model ReleasesDGX agent

arXiv:2605.07186v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) benchmarks primarily focus on syntactically correct inputs, leaving a significant gap in evaluation on imperfect t

The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks

Model ReleasesDGX agent

arXiv:2605.07093v1 Announce Type: cross Abstract: The Translation Tax is often treated as a scalar: translated benchmarks are assumed to inflate scores by preserving English-source cues. We audit this

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model

Model ReleasesDGX agent

arXiv:2602.04774v2 Announce Type: replace-cross Abstract: Setting the learning rate (LR) for a deep learning model is a critical part of successful training. Choosing LRs is often done empirically wit

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Model ReleasesDGX agent

arXiv:2601.23143v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

Model ReleasesDGX agent

arXiv:2510.01290v2 Announce Type: replace Abstract: The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key-value (

Today we’re launching the OpenAI Deployment Company to help businesses build and deploy AI. It's majority-owned and controlled by OpenAI. It…

Model ReleasesDGX agent

Today we’re launching the OpenAI Deployment Company to help businesses build and deploy AI. It's majority-owned and controlled by OpenAI. It brings together 19 leading investment firms, consultancies,

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models

Model ReleasesDGX agent

arXiv:2605.06683v1 Announce Type: cross Abstract: Transformer-based large language models are in some respects limited by the quadratic time and space computational complexity of attention. We introdu

Tool Calling is Linearly Readable and Steerable in Language Models

Model ReleasesDGX agent

arXiv:2605.07990v1 Announce Type: cross Abstract: When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. Probing 12 ins

Tools as Continuous Flow for Evolving Agentic Reasoning

Model ReleasesDGX agent

arXiv:2605.07339v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a s

Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings

Model ReleasesDGX agent

arXiv:2605.07158v1 Announce Type: cross Abstract: Vector search and retrieval-augmented generation (RAG) rest on the assumption that cosine similarity between text embeddings reflects conceptual relat

Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion

Model ReleasesDGX agent

arXiv:2605.07013v1 Announce Type: new Abstract: Diffusion language models (DLMs) promise parallel, order-agnostic generation, but on standard benchmarks they have historically lagged behind autoregres

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

Model ReleasesDGX agent

arXiv:2605.07593v1 Announce Type: new Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams,

Tracing Uncertainty in Language Model 'Reasoning'

Model ReleasesDGX agent

arXiv:2605.07776v1 Announce Type: cross Abstract: Language model (LM) 'reasoning', commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics u

Traffic Scenario Orchestration from Language via Constraint Satisfaction

Model ReleasesDGX agent

arXiv:2605.06966v1 Announce Type: new Abstract: Autonomous vehicles (AVs) require extensive testing in simulation, but test case generation for driving scenarios is laborious. The desired scenarios ar

Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers

Model ReleasesDGX agent

arXiv:2605.07772v1 Announce Type: new Abstract: Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically,

Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

Model ReleasesDGX agent

arXiv:2605.07924v1 Announce Type: cross Abstract: Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Dis

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

Model ReleasesDGX agent

arXiv:2601.18744v2 Announce Type: replace Abstract: Time series are ubiquitous in real-world scenarios and crucial for applications ranging from energy management to traffic control. Consequently, the

UNCOM: Zero-shot Context-Aware Command Understanding for Tabletop Scenarios

Model ReleasesDGX agent

arXiv:2410.06355v3 Announce Type: replace-cross Abstract: This paper presents UNCOM, a novel hybrid framework for interpreting natural human commands in tabletop scenarios. The system integrates multi

Understanding Robustness of Model Editing in Code LLMs

Model ReleasesDGX agent

arXiv:2511.03182v2 Announce Type: replace-cross Abstract: Large language models (LLMs) for code are increasingly used in software development, but they remain static after pretraining while APIs and s

Uneven Evolution of Cognition Across Generations of Generative AI Models

Model ReleasesDGX agent

arXiv:2605.06815v1 Announce Type: new Abstract: The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task perfor

Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts

Model ReleasesDGX agent

arXiv:2605.07395v1 Announce Type: cross Abstract: Efficient routing across multiple LLMs enables cost-quality tradeoffs by directing queries to the cheapest capable model. Prior work attributes routin

Using LLM in the shebang line of a script

Model ReleasesDGX agent

TIL: Using LLM in the shebang line of a script Kim_Bruning on Hacker News: But seriously, you can put a shebang on an english text file now (if you're sufficiently brave) [...] This inspired me to loo

Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset

Model ReleasesDGX agent

arXiv:2602.16571v2 Announce Type: replace Abstract: Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrie

v0.30.0-rc14: Merge remote-tracking branch 'upstream/main' into llama-runner-phase-0

Model ReleasesDGX agent

This is a release candidate version of Ollama that merges updates from the main development branch into the llama-runner-phase-0 branch, likely incorporating bug fixes and features in preparation for

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

Model ReleasesDGX agent

arXiv:2605.07872v1 Announce Type: cross Abstract: Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely l

Visual Text Compression as Measure Transport

Model ReleasesDGX agent

arXiv:2605.06708v1 Announce Type: cross Abstract: Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language mod

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts

Model ReleasesDGX agent

arXiv:2605.06175v2 Announce Type: replace Abstract: Vision-language-action (VLA) models inherit rich visual-semantic priors from pre-trained vision-language backbones, but adapting them to robotic con

📣We're calling for ambassadors! Whether you're a developer with great technical taste or a local community leader who loves bringing people…

Model ReleasesDGX agent

📣We're calling for ambassadors! Whether you're a developer with great technical taste or a local community leader who loves bringing people together, we'd love to have you join us. Visit the website b

We’ve also agreed to acquire Tomoro, which will bring 150 experienced Forward Deployed Engineers and Deployment Specialists to the OpenAI De…

Model ReleasesDGX agent

OpenAI announced the acquisition of Tomoro, a company that will add 150 experienced Forward Deployed Engineers and Deployment Specialists to OpenAI's team. This acquisition appears focused on expandin

When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

Model ReleasesDGX agent

arXiv:2605.07260v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models route each token to a small subset of experts, but whether the routes selected by a trained top-k router are

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

Model ReleasesDGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents

Model ReleasesDGX agent

arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

Model ReleasesDGX agent

arXiv:2605.07114v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language model

Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions

Model ReleasesDGX agent

arXiv:2605.07984v1 Announce Type: cross Abstract: We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forwar

Why Self-Inconsistency Arises in GNN Explanations and How to Exploit It

Model ReleasesDGX agent

arXiv:2605.07527v1 Announce Type: cross Abstract: Recent work has observed that explanations produced by Self-Interpretable Graph Neural Networks (SI-GNNs) can be self-inconsistent: when the model is

WiCER: Wiki-memory Compile, Evaluate, Refine Iterative Knowledge Compilation for LLM Wiki Systems

Model ReleasesDGX agent

arXiv:2605.07068v1 Announce Type: cross Abstract: The LLM Wiki pattern, to compile and provide domain knowledge into a persistent artifact and serve it to LLMs via KV cache inference, promises context

Worth reading

Model ReleasesDGX agent

This post from Elon Musk's X account likely recommends or highlights content that Musk finds valuable or important, though without access to the specific tweet content, the exact subject matter cannot

would you call it a superapp?

Model ReleasesDGX agent

would you call it a superapp? After being a Claude Code devotee for a year, I finally tried Codex on a new project this weekend. Once again, in the matter of a few months, it feels like the world chan

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

Model ReleasesDGX agent

arXiv:2605.07579v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing ap

← Previous
1…274275276277278…376
Next →