AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,555 results
7 May 2026

MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2605.03675v1 Announce Type: new Abstract: Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over

MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

Model ReleasesDGX agent

arXiv:2605.03952v1 Announce Type: cross Abstract: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The chal

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.04058v1 Announce Type: new Abstract: Parameter-efficient transfer learning (PETL) has emerged as a pivotal paradigm for adapting pre-trained foundation models to downstream tasks, significa

MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge

Model ReleasesDGX agent

arXiv:2605.05175v1 Announce Type: cross Abstract: Background: Existing MRI LLM benchmarks rely mainly on review-book multiple-choice questions, where top proprietary models already score highly, limit

Multi-Token Prediction (MTP) for LLaMA.cpp! Running Gemma4 local model 1.5x faster. We patched LLaMA.cpp. Quantized Gemma 4 assistant models…

Model ReleasesDGX agent

Multi-Token Prediction (MTP) for LLaMA.cpp! Running Gemma4 local model 1.5x faster. We patched LLaMA.cpp. Quantized Gemma 4 assistant models into GGUF format. We ran tests on a MacBook Pro M5Max. Gemm

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains

Model ReleasesDGX agent

arXiv:2511.06452v3 Announce Type: replace Abstract: Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Curren

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations…

Model ReleasesDGX agent

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read

NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise

Model ReleasesDGX agent

arXiv:2605.04313v1 Announce Type: new Abstract: Causal reasoning in natural language requires identifying relevant variables, understanding their interactions, and reasoning about effects and interven

Not All That Is Fluent Is Factual: Investigating Hallucinations of Large Language Models in Academic Writing

Model ReleasesDGX agent

arXiv:2605.04171v1 Announce Type: new Abstract: Large Language models (LLMs) show extraordinary abilities, but they are still prone to hallucinations, especially when we use them for generating Academ

Notes on the xAI/Anthropic data center deal

Model ReleasesDGX agent

There weren't a lot of big new announcements from Anthropic at yesterday's Code w/ Claude event, but the biggest by far was the deal they've struck with SpaceX/xAI to use 'all of the capacity of their

Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages

Model ReleasesDGX agent

arXiv:2605.04208v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive multilingual capabilities for well-resourced languages, yet their performance on low-resource

Open Models x Headless Agent Execution 🔥

Model ReleasesDGX agent

Open Models x Headless Agent Execution 🔥 your daily reminder that open models are plenty capable for a lot of coding work. easiest place to feel that out is deepagents! swap the model and go. i've bee

Open-Source Image Editing Models Are Zero-Shot Vision Learners

Model ReleasesDGX agent

arXiv:2605.04566v1 Announce Type: cross Abstract: Recent studies have shown that large generative models can solve vision tasks they were not explicitly trained for. However, existing evidence relies

OpenAI for Excel is quite useful (as is Claude for Excel), so it is surprising, that, unlike Claude, there is no OpenAI for PowerPoint, espe…

Model ReleasesDGX agent

OpenAI for Excel is quite useful (as is Claude for Excel), so it is surprising, that, unlike Claude, there is no OpenAI for PowerPoint, especially because it is where OpenAI has a big advantage: Image

OpenClaw and Claude can put your AI-generated podcasts in Spotify

Model ReleasesDGX agent

Save to Spotify is a new command-line tool designed specifically for AI agents like OpenClaw, Claude Code, or OpenAI Codex. If you're the kind of person who collects research on a topic, then feeds it

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

Model ReleasesDGX agent

arXiv:2601.22725v3 Announce Type: replace Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remain

Optimal Control with Natural Images: Efficient Reinforcement Learning using Overcomplete Sparse Codes

Model ReleasesDGX agent

arXiv:2412.08893v3 Announce Type: replace Abstract: Optimal control and sequential decision making are widely used in many complex tasks. Optimal control over a sequence of natural images is a first s

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

Model ReleasesDGX agent

arXiv:2605.04738v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities. However, their massive parameter scale leads to significant resource consumption

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs

Model ReleasesDGX agent

arXiv:2605.04665v1 Announce Type: new Abstract: When the substantive content of a request is rewritten, do large language models still answer in the format the original task asked for? We find that th

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

Model ReleasesDGX agent

arXiv:2509.24943v2 Announce Type: replace Abstract: Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Althou

Physics-Grounded Multi-Agent Architecture for Traceable, Risk-Aware Human-AI Decision Support in Manufacturing

Model ReleasesDGX agent

arXiv:2605.04003v1 Announce Type: cross Abstract: High-precision CNC machining of free-form aerospace components requires bounded compensations informed by inspection, simulation, and process knowledg

Privacy-Preserving Empathy Detection in Video Interactions

Model ReleasesDGX agent

arXiv:2504.10808v3 Announce Type: replace Abstract: Detecting empathy from video interactions has emerging applications, yet raw videos that could be used for training AI models are rarely available d

Probing Structural Mathematical Reasoning in Language Models with Algebraic Trapdoors

Model ReleasesDGX agent

arXiv:2605.04352v1 Announce Type: new Abstract: We introduce a benchmark suite for evaluating structural mathematical reasoning in language models, built on subgroup-construction problems in SL(3, Z)

Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification

Model ReleasesDGX agent

arXiv:2605.05027v1 Announce Type: new Abstract: Lifelong person re-identification (LReID) aims to train a generalizable model with sequentially collected data. However, such models often suffer from s

Provable Non-Convex Euclidean Distance Matrix Completion: Geometry, Reconstruction, and Robustness

Model ReleasesDGX agent

arXiv:2508.00091v3 Announce Type: replace-cross Abstract: The problem of recovering the configuration of points from their partial pairwise distances, referred to as the Euclidean Distance Matrix Comp

PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation

Model ReleasesDGX agent

arXiv:2605.05159v1 Announce Type: new Abstract: We present our system for SemEval-2026 Task 9: Multilingual Polarization Detection, a binary classification task spanning 22 languages. Our approach fin

QKVShare: Quantized KV-Cache Handoff for Multi-Agent On-Device LLMs

Model ReleasesDGX agent

arXiv:2605.03884v1 Announce Type: new Abstract: Multi-agent LLM systems on edge devices need to hand off latent context efficiently, but the practical choices today are expensive re-prefill or full-pr

Quantum-inspired Reinforcement Learning for Synthesizable Drug Design

Model ReleasesDGX agent

arXiv:2409.09183v2 Announce Type: replace Abstract: Synthesizable molecular design (also known as synthesizable molecular optimization) is a fundamental problem in drug discovery, and involves designi

Real-Time Evaluation of Autonomous Systems under Adversarial Attacks

Model ReleasesDGX agent

arXiv:2605.03491v1 Announce Type: new Abstract: Most evaluations of autonomous driving policies under adversarial conditions are conducted in simulation, due to cost efficiency and the absence of phys

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

Model ReleasesDGX agent

arXiv:2605.03361v2 Announce Type: new Abstract: As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organizati

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

Model ReleasesDGX agent

arXiv:2605.04019v1 Announce Type: new Abstract: AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a

Regime-Conditioned Evaluation in Multi-Context Bayesian Optimization

Model ReleasesDGX agent

arXiv:2605.04895v1 Announce Type: new Abstract: Published transfer-BO comparisons often estimate an average treatment effect of acquisition choice over hidden regime variables, while practitioners nee

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.03426v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have broad potential in privacy-sensitive domains such as healthcare and finance, yet strict data-sharing constraints rend

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

Model ReleasesDGX agent

arXiv:2605.04539v1 Announce Type: new Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference si

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos

Model ReleasesDGX agent

arXiv:2412.03077v2 Announce Type: replace Abstract: 4D reconstruction from casually captured monocular videos is challenging due to inherent ambiguity in reconstructing dynamic 3D geometry. To address

Saw this and thought 'yes! ChatGPT voice mode is going to stop acting like a two-year-model' but that upgrade hasn't shipped just yet

Model ReleasesDGX agent

Saw this and thought 'yes! ChatGPT voice mode is going to stop acting like a two-year-model' but that upgrade hasn't shipped just yet Introducing GPT-Realtime-2 in the API: our most intelligent voice

Scalable Object Detection in the Car Interior With Vision Foundation Models

Model ReleasesDGX agent

arXiv:2508.19651v2 Announce Type: replace Abstract: AI tasks in the car interior like identifying and localizing externally introduced objects is crucial for response quality of personal assistants. H

Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber

Model ReleasesDGX agent

This document likely describes OpenAI's implementation of trusted access controls and security features in GPT-5.5 and a specialized GPT-5.5-Cyber variant designed for cybersecurity applications. It p

Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics

Model ReleasesDGX agent

arXiv:2605.04893v1 Announce Type: cross Abstract: Large language models hallucinate in predictable ways: attention routing fails by over-concentrating on a narrow set of positions, or by spreading so

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

Model ReleasesDGX agent

arXiv:2605.04221v1 Announce Type: new Abstract: Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often pri

Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval

Model ReleasesDGX agent

arXiv:2605.05189v1 Announce Type: cross Abstract: How many key-value associations can a dimes d linear memory store? We show that the answer depends not only on the d^2 degrees of freedom in the memor

“She said the theme of this party is the industrial age. And you came in dressed like a train wreck.” Asking AIs to think of the equivalent …

Model ReleasesDGX agent

“She said the theme of this party is the industrial age. And you came in dressed like a train wreck.” Asking AIs to think of the equivalent to this Hold Steady lyric, but for AI. Claude was the clear

Single-Position Intervention Fails: Distributed Output Templates Drive In-Context Learning

Model ReleasesDGX agent

arXiv:2605.04061v1 Announce Type: cross Abstract: Understanding how large language models encode task identity from few-shot demonstrations is a central open problem in mechanistic interpretability. P

SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents

Model ReleasesDGX agent

arXiv:2605.03353v1 Announce Type: cross Abstract: LLM-Agents have evolved into autonomous systems for complex task execution, with the SKILL.md specification emerging as a de facto standard for encaps

Skill Neologisms: Towards Skill-based Continual Learning

Model ReleasesDGX agent

arXiv:2605.04970v1 Announce Type: new Abstract: Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

Model ReleasesDGX agent

arXiv:2511.06754v3 Announce Type: replace-cross Abstract: Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation rep

Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction

Model ReleasesDGX agent

arXiv:2605.04072v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have been applied to large language models and protein language models, but not systematically to electronic health record

SpecPL: Disentangling Spectral Granularity for Prompt Learning

Model ReleasesDGX agent

arXiv:2605.04504v1 Announce Type: cross Abstract: Existing prompt learning for VLMs exhibits a modality asymmetry, predominantly optimizing text tokens while still relying on frozen visual encoder as

Spotify launches Save to Spotify, a command-line tool that allows AI agents to upload AI-generated audio summaries and personal podcasts to a user's account (Terrence O'Brien/The Verge)

Model ReleasesDGX agent

Terrence O'Brien / The Verge: Spotify launches Save to Spotify, a command-line tool that allows AI agents to upload AI-generated audio summaries and personal podcasts to a user's account — A new comma

Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense

Model ReleasesDGX agent

arXiv:2605.03034v1 Announce Type: new Abstract: Agentic systems involved in high-stake decision-making under adversarial pressure need formal guarantees not offered by existing approaches. Motivated b

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

Model ReleasesDGX agent

arXiv:2605.04453v1 Announce Type: new Abstract: In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetic

Stage Light is Sequence^2: Multi-Light Control via Imitation Learning

Model ReleasesDGX agent

arXiv:2605.03660v1 Announce Type: cross Abstract: Music-inspired Automatic Stage Lighting Control (ASLC) has gained increasing attention in recent years due to the substantial time and financial costs

Storage Is Not Memory: A Retrieval-Centered Architecture for Agent Recall

Model ReleasesDGX agent

arXiv:2605.04897v1 Announce Type: new Abstract: Extraction at ingestion is the wrong primitive for agent memory: content discarded before the query is known cannot be recovered at retrieval time. We p

StoryAlign: Evaluating and Training Reward Models for Story Generation

Model ReleasesDGX agent

arXiv:2605.04831v1 Announce Type: new Abstract: Story generation aims to automatically produce coherent, structured, and engaging narratives. Although large language models (LLMs) have significantly a

SWAN: Semantic Watermarking with Abstract Meaning Representation

Model ReleasesDGX agent

arXiv:2605.04305v1 Announce Type: new Abstract: We introduce SWAN (Semantic Watermarking with Abstract Meaning Representation), a novel framework that embeds watermark signatures into the semantic str

Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors

Model ReleasesDGX agent

arXiv:2602.00305v2 Announce Type: replace-cross Abstract: LLM-based vulnerability detectors are increasingly deployed in CI/CD security gating, yet their resilience to evasion under syntax- and compil

TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding

Model ReleasesDGX agent

arXiv:2605.04962v1 Announce Type: new Abstract: Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular dat

TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference

Model ReleasesDGX agent

arXiv:2603.26498v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) power platforms like ChatGPT, Gemini, and Copilot, enabling richer interactions with text, images, an

Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

Model ReleasesDGX agent

arXiv:2605.04426v1 Announce Type: new Abstract: We introduce Telegraph English (TE), a prompt-compression protocol that rewrites natural language into a symbol-rich, formally-structured dialect. Where

Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?

Model ReleasesDGX agent

arXiv:2605.03195v1 Announce Type: new Abstract: Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow responsibiliti

← Previous
1…278279280281282…376
Next →