AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,574 results
15 May 2026

Do Coding Agents Understand Least-Privilege Authorization?

Model ReleasesDGX agent

arXiv:2605.14859v1 Announce Type: cross Abstract: As coding agents gain access to shells, repositories, and user files, least-privilege authorization becomes a prerequisite for safe deployment: an age

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Model ReleasesDGX agent

arXiv:2512.13609v2 Announce Type: replace Abstract: We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transf

'do you want meet in the Devin train car?' 'oh actually I'm in the Langsmith car near the front' 'Let's just meet in front of the Mistral wa…

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

This appears to be a humorous exchange from X/Twitter creator Harrison Chase referencing train cars named after AI/ML tools and services (Devin, Langsmith, and Mistral), likely making a joke about the

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

Model ReleasesDGX agent

arXiv:2605.14473v1 Announce Type: cross Abstract: The Context-Compliance Regime in Retrieval-Augmented Generation (RAG) occurs when retrieved context dominates the final answer even when it conflicts

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

Model ReleasesDGX agent

arXiv:2605.14420v1 Announce Type: new Abstract: Current Large Language Models (LLMs) typically rely on coarse-grained national labels for pluralistic value alignment. However, such macro-level supervi

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

Model ReleasesDGX agent

arXiv:2605.14842v1 Announce Type: new Abstract: Humans naturally communicate through abstract concepts like 'mood'. However, current image editing benchmarks focus primarily on explicit, literal comma

Efficient Online Conformal Selection with Limited Feedback

Model ReleasesDGX agent

arXiv:2605.14953v1 Announce Type: new Abstract: We address the problem of conformal selection, where an agent must select a minimal subset of options to ensure that at least one ``success'' is identif

Elastic Spiking Transformers for Efficient Gesture Understanding

Model ReleasesDGX agent

arXiv:2605.13869v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs), particularly Spiking Transformers, offer energy-efficient processing of event-based sensor data for healthcare applica

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring

Model ReleasesDGX agent

arXiv:2605.14589v1 Announce Type: new Abstract: Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and comp

Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ABHFA-Net

Model ReleasesDGX agent

arXiv:2510.18326v3 Announce Type: replace Abstract: The rising incidence of natural and human-induced disasters necessitates robust visual recognition systems capable of operating under limited labele

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

Model ReleasesDGX agent

arXiv:2605.15199v1 Announce Type: cross Abstract: Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and location

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

Model ReleasesDGX agent

arXiv:2605.15179v1 Announce Type: cross Abstract: Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of d

EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

Model ReleasesDGX agent

arXiv:2605.13941v1 Announce Type: cross Abstract: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixe

Exemplar Partitioning for Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2605.14347v1 Announce Type: new Abstract: We introduce Exemplar Partitioning (EP), an unsupervised method for constructing interpretable feature dictionaries from large language model activation

ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Model ReleasesDGX agent

arXiv:2605.14153v1 Announce Type: cross Abstract: Exploitation is not a binary event. It is a ladder of acquiring progressive capabilities, from executing a single buggy line of code to taking full co

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study

Model ReleasesDGX agent

arXiv:2605.14845v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigor

Fast Rates for Inverse Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.14599v1 Announce Type: cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) with linear reward

FedStain: Modeling Higher-Order Stain Statistics for Federated Domain Generalization in Computational Pathology

Model ReleasesDGX agent

arXiv:2605.14590v1 Announce Type: new Abstract: Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Dom

FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

Model ReleasesDGX agent

arXiv:2604.06757v2 Announce Type: replace Abstract: Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We chal

Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution

Model ReleasesDGX agent

arXiv:2605.15138v1 Announce Type: cross Abstract: Standard unlearning evaluations measure behavioral suppression in full precision, immediately after training, despite every deployed language model be

From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents

Model ReleasesDGX agent

arXiv:2605.14034v1 Announce Type: new Abstract: Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cogni

From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG

Model ReleasesDGX agent

arXiv:2605.15019v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fi

From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement

Model ReleasesDGX agent

arXiv:2605.14912v1 Announce Type: new Abstract: Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or prop

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

Model ReleasesDGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

Fusion-fission forecasts when AI will shift to undesirable behavior

Model ReleasesDGX agent

arXiv:2605.14218v1 Announce Type: new Abstract: The key problem facing ChatGPT-like AI's use across society is that its behavior can shift, unnoticed, from desirable to undesirable -- encouraging self

FutureSim: Replaying World Events to Evaluate Adaptive Agents

Model ReleasesDGX agent

arXiv:2605.15188v1 Announce Type: cross Abstract: AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently m

G-SHARP: Gaussian Surgical Hardware Accelerated Real-time Pipeline

Model ReleasesDGX agent

arXiv:2512.02482v2 Announce Type: replace Abstract: We propose G-SHARP, a commercially compatible, real-time surgical scene reconstruction framework designed for minimally invasive procedures that req

Gemini Live Agent Challenge: Announcing the winners and highlights

Model ReleasesDGX agent

The Gemini Live Agent Challenge is officially in the books! We challenged developers worldwide to break out of the traditional 'text box' paradigm by building next-generation AI agents. From our initi

Gemma-4-31B-it-Pearl supports 256K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powere…

Model ReleasesDGX agent

Gemma-4-31B-it-Pearl supports 256K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powered endpoint. Eventually, we plan to expand our Pearl powered

Gemma-4-31B-it-Pearl supports text and image input, 256K context, configurable thinking, function calling, and JSON mode. This is Together A…

Model ReleasesDGX agent

Gemma-4-31B-it-Pearl supports text and image input, 256K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powered endpoint. Eventually, we plan to exp

GenCircuit-RL: Reinforcement Learning from Hierarchical Verification for Genetic Circuit Design

Model ReleasesDGX agent

arXiv:2605.14215v1 Announce Type: new Abstract: Genetic circuit design remains a laborious, expert-driven process despite decades of progress in synthetic biology. We study this problem through code g

GenExam: A Multidisciplinary Text-to-Image Exam

Model ReleasesDGX agent

arXiv:2509.14232v5 Announce Type: replace Abstract: Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style bench

GFMate: Empowering Graph Foundation Models with Test-time Prompt Tuning

Model ReleasesDGX agent

arXiv:2605.14809v1 Announce Type: new Abstract: Graph prompt tuning has shown great potential in graph learning by introducing trainable prompts to enhance the model performance in conventional single

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

Model ReleasesDGX agent

arXiv:2602.06718v2 Announce Type: replace-cross Abstract: Citations provide the basis for trusting scientific claims; when they are invalid or fabricated, this trust collapses. With the advent of Larg

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

Model ReleasesDGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

Model ReleasesDGX agent

arXiv:2605.14841v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) has become the dominant paradigm for parameter-efficient fine-tuning (PEFT) of large language models (LLMs). However, its b

GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration

Model ReleasesDGX agent

arXiv:2605.13848v1 Announce Type: new Abstract: Agentic LLM frameworks that rely on prompted orchestration, where the model itself determines workflow transitions, often suffer from hallucinated routi

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Model ReleasesDGX agent

arXiv:2605.14498v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems t

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

Model ReleasesDGX agent

arXiv:2605.14821v1 Announce Type: new Abstract: Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit

Herculean: An Agentic Benchmark for Financial Intelligence

Model ReleasesDGX agent

arXiv:2605.14355v1 Announce Type: new Abstract: As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carr

Hidden State Poisoning Attacks against Mamba-based Language Models

Model ReleasesDGX agent

arXiv:2601.01972v4 Announce Type: cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. Yet, their advers

HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

Model ReleasesDGX agent

arXiv:2605.15024v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images

Holistic Evaluation and Failure Diagnosis of AI Agents

Model ReleasesDGX agent

arXiv:2605.14865v1 Announce Type: new Abstract: AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, an

How business operations teams use Codex

Model ReleasesDGX agent

This document from OpenAI describes how business operations teams leverage Codex, OpenAI's code generation model, to automate and streamline their workflows. It likely covers practical applications su

How data science teams use Codex

Model ReleasesDGX agent

This OpenAI Academy resource explores practical applications of Codex, their AI code generation model, within data science workflows and teams. It likely covers how data scientists leverage Codex to a

How sales teams use Codex

Model ReleasesDGX agent

This resource from OpenAI's Academy demonstrates practical applications of Codex, their code-generation AI model, within sales team workflows. It likely covers how sales professionals can leverage Cod

How Sensitive Are Radiomic AI Models to Acquisition Parameters?

Model ReleasesDGX agent

arXiv:2605.14667v1 Announce Type: new Abstract: A main barrier for the deployment of AI radiomic systems in clinical routine is their drop in performance under heterogeneous multicentre acquisition pr

Hyperspectral Image Land Cover Captioning Dataset for Vision Language Models

Model ReleasesDGX agent

arXiv:2505.12217v2 Announce Type: replace Abstract: We introduce HyperCap, the first large-scale hyperspectral captioning dataset designed to enhance model performance and effectiveness in remote sens

i appreciate how seriously the team always takes these reports (even when the answer turns out to be 'i got used to the current level of mag…

Model ReleasesDGX agent

i appreciate how seriously the team always takes these reports (even when the answer turns out to be 'i got used to the current level of magic and now i'd like more please') Codex team is aware of rep

In April of 2021 horrible, racist graffiti was found in the stairwell at Albion College in Michigan. Sprawled on the walls were hateful mess…

Model ReleasesDGX agent

In April of 2021 horrible, racist graffiti was found in the stairwell at Albion College in Michigan. Sprawled on the walls were hateful messages like 'KKK White Power' & 'Die N*gger Please' Liberals c

In-Context Learning for Data-Driven Censored Inventory Control

Model ReleasesDGX agent

arXiv:2605.14840v1 Announce Type: new Abstract: We study inventory control with decision-dependent censoring, focusing on the censored or repeated newsvendor (R-NV), where each order quantity determin

Indian Wedding System Optimization (IWSO): A Novel Socially Inspired Metaheuristic with Operational Design and Analysis

Model ReleasesDGX agent

arXiv:2605.13871v1 Announce Type: cross Abstract: This paper presents a novel population-based metaheuristic, Indian Wedding System Optimization (IWSO), inspired by the socio-cultural dynamics of trad

Inference is becoming the largest compute market and energy consumer in AI. Pearl turns inference CapEx of hyperscalers into a profit center…

Model ReleasesDGX agent

Inference is becoming the largest compute market and energy consumer in AI. Pearl turns inference CapEx of hyperscalers into a profit center: every LLM token produced by GPUs can simultaneously genera

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Model ReleasesDGX agent

arXiv:2605.14712v1 Announce Type: cross Abstract: Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators

Introducing Gemma-4-31B-it-Pearl on Together AI, Pearl Research Labs’ instruction-tuned checkpoint of Gemma 4 31B powered by @prlnet Proof o…

Model ReleasesDGX agent

Introducing Gemma-4-31B-it-Pearl on Together AI, Pearl Research Labs’ instruction-tuned checkpoint of Gemma 4 31B powered by @prlnet Proof of Useful Work protocol. AI natives can now use this Pearl mo

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.13851v1 Announce Type: new Abstract: Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI d

IPR-1: Interactive Physical Reasoner

Model ReleasesDGX agent

arXiv:2511.15407v3 Announce Type: replace Abstract: Humans learn by observing, interacting with environments, and internalizing physics and causality. Here, we aim to ask whether an agent can similarl

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

Model ReleasesDGX agent

arXiv:2605.15184v1 Announce Type: new Abstract: Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools,

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2512.12772v2 Announce Type: replace-cross Abstract: Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models

K-Models: a Flexible and Interpretable Method for Ordinal Clustering with Application to Antigen-Antibody Interaction Profiles

Model ReleasesDGX agent

arXiv:2605.14828v1 Announce Type: cross Abstract: Existing clustering methods for functional data often prioritize partitioning accuracy over interpretability, making it challenging to extract meaning

← Previous
1…247248249250251…377
Next →