AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
15 May 2026

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

Model ReleasesDGX agent

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models

Model ReleasesDGX agent

arXiv:2605.14897v1 Announce Type: cross Abstract: Despite many successful attempts at explaining Deep Reinforcement Learning policies using distillation, it remains difficult to balance the performanc

CUICurate: A GraphRAG-based Framework for Automated Clinical Concept Curation for NLP applications

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2602.17949v2 Announce Type: replace-cross Abstract: Background: Clinical named entity recognition tools commonly map free text to Unified Medical Language System (UMLS) Concept Unique Identifier

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves

Model ReleasesDGX agent

arXiv:2605.14068v1 Announce Type: new Abstract: We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of extbf{756 images} of pairwise non-

Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning

Model ReleasesDGX agent

arXiv:2605.14386v1 Announce Type: cross Abstract: We present Darwin Family, a framework for training-free evolutionary merging of large language models via gradient-free weight-space recombination. We

Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

Model ReleasesDGX agent

arXiv:2605.14379v1 Announce Type: cross Abstract: Finding approximate equilibria for large-scale imperfect-information competitive games such as StarCraft, Dota, and CounterStrike remains computationa

Databricks brings GPT-5.5 to enterprise agent workflows

Model ReleasesDGX agent

Databricks has integrated OpenAI's GPT-5.5 model into enterprise agent workflows, enabling organizations to build and deploy AI agents with advanced language capabilities. This partnership leverages D

Dataiku launches governed AI workflow builder inside Snowflake

Model ReleasesDGX agent

Dataiku Inc. today announced a deeper integration with Snowflake Inc. aimed at simplifying the creation of enterprise AI agents while preserving governance and operational controls, a growing concern

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia

Model ReleasesDGX agent

arXiv:2509.23023v3 Announce Type: replace Abstract: Large language models are increasingly deployed in multi-agent settings whose outcomes hinge on social intelligence, motivating evaluations of their

Deep Image Segmentation via Discriminant Feature Learning

Model ReleasesDGX agent

arXiv:2605.14609v1 Announce Type: new Abstract: Accurate image segmentation remains challenging, particularly in generating sharp, confident boundaries. While modern architectures have advanced the fi

Denoising-GS: Gaussian Splatting with Spatial-aware Denoising

Model ReleasesDGX agent

arXiv:2605.14880v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable success in high-fidelity Novel View Synthesis (NVS), yet the optimization proce

Derivation Prompting: A Logic-Based Method for Improving Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.14053v1 Announce Type: cross Abstract: The application of Large Language Models to Question Answering has shown great promise, but important challenges such as hallucinations and erroneous

Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)

Model ReleasesDGX agent

arXiv:2511.13397v2 Announce Type: replace-cross Abstract: The remarkable progress of Vision-Language Models (VLMs) on a variety of tasks has raised interest in their application to automated driving.

Discovering Physical Directions in Weight Space: Composing Neural PDE Experts

Model ReleasesDGX agent

arXiv:2605.14546v1 Announce Type: new Abstract: Recent advances in neural operators have made partial differential equation (PDE) surrogate modeling increasingly scalable and transferable through larg

Distribution-Aware Algorithm Design with LLM Agents

Model ReleasesDGX agent

arXiv:2605.14141v1 Announce Type: new Abstract: We study learning when the learned object is executable solver code rather than a predictor. In this setting, correctness is not enough: two solvers may

Do Coding Agents Understand Least-Privilege Authorization?

Model ReleasesDGX agent

arXiv:2605.14859v1 Announce Type: cross Abstract: As coding agents gain access to shells, repositories, and user files, least-privilege authorization becomes a prerequisite for safe deployment: an age

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Model ReleasesDGX agent

arXiv:2512.13609v2 Announce Type: replace Abstract: We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transf

'do you want meet in the Devin train car?' 'oh actually I'm in the Langsmith car near the front' 'Let's just meet in front of the Mistral wa…

Model ReleasesDGX agent

This appears to be a humorous exchange from X/Twitter creator Harrison Chase referencing train cars named after AI/ML tools and services (Devin, Langsmith, and Mistral), likely making a joke about the

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

Model ReleasesDGX agent

arXiv:2605.14473v1 Announce Type: cross Abstract: The Context-Compliance Regime in Retrieval-Augmented Generation (RAG) occurs when retrieved context dominates the final answer even when it conflicts

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

Model ReleasesDGX agent

arXiv:2605.14420v1 Announce Type: new Abstract: Current Large Language Models (LLMs) typically rely on coarse-grained national labels for pluralistic value alignment. However, such macro-level supervi

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

Model ReleasesDGX agent

arXiv:2605.14842v1 Announce Type: new Abstract: Humans naturally communicate through abstract concepts like 'mood'. However, current image editing benchmarks focus primarily on explicit, literal comma

Efficient Online Conformal Selection with Limited Feedback

Model ReleasesDGX agent

arXiv:2605.14953v1 Announce Type: new Abstract: We address the problem of conformal selection, where an agent must select a minimal subset of options to ensure that at least one ``success'' is identif

Elastic Spiking Transformers for Efficient Gesture Understanding

Model ReleasesDGX agent

arXiv:2605.13869v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs), particularly Spiking Transformers, offer energy-efficient processing of event-based sensor data for healthcare applica

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring

Model ReleasesDGX agent

arXiv:2605.14589v1 Announce Type: new Abstract: Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and comp

Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ABHFA-Net

Model ReleasesDGX agent

arXiv:2510.18326v3 Announce Type: replace Abstract: The rising incidence of natural and human-induced disasters necessitates robust visual recognition systems capable of operating under limited labele

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

Model ReleasesDGX agent

arXiv:2605.15199v1 Announce Type: cross Abstract: Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and location

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

Model ReleasesDGX agent

arXiv:2605.15179v1 Announce Type: cross Abstract: Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of d

EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

Model ReleasesDGX agent

arXiv:2605.13941v1 Announce Type: cross Abstract: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixe

Exemplar Partitioning for Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2605.14347v1 Announce Type: new Abstract: We introduce Exemplar Partitioning (EP), an unsupervised method for constructing interpretable feature dictionaries from large language model activation

ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Model ReleasesDGX agent

arXiv:2605.14153v1 Announce Type: cross Abstract: Exploitation is not a binary event. It is a ladder of acquiring progressive capabilities, from executing a single buggy line of code to taking full co

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study

Model ReleasesDGX agent

arXiv:2605.14845v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigor

Fast Rates for Inverse Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.14599v1 Announce Type: cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) with linear reward

FedStain: Modeling Higher-Order Stain Statistics for Federated Domain Generalization in Computational Pathology

Model ReleasesDGX agent

arXiv:2605.14590v1 Announce Type: new Abstract: Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Dom

FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

Model ReleasesDGX agent

arXiv:2604.06757v2 Announce Type: replace Abstract: Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We chal

Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution

Model ReleasesDGX agent

arXiv:2605.15138v1 Announce Type: cross Abstract: Standard unlearning evaluations measure behavioral suppression in full precision, immediately after training, despite every deployed language model be

From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents

Model ReleasesDGX agent

arXiv:2605.14034v1 Announce Type: new Abstract: Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cogni

From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG

Model ReleasesDGX agent

arXiv:2605.15019v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fi

From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement

Model ReleasesDGX agent

arXiv:2605.14912v1 Announce Type: new Abstract: Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or prop

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

Model ReleasesDGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

Fusion-fission forecasts when AI will shift to undesirable behavior

Model ReleasesDGX agent

arXiv:2605.14218v1 Announce Type: new Abstract: The key problem facing ChatGPT-like AI's use across society is that its behavior can shift, unnoticed, from desirable to undesirable -- encouraging self

FutureSim: Replaying World Events to Evaluate Adaptive Agents

Model ReleasesDGX agent

arXiv:2605.15188v1 Announce Type: cross Abstract: AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently m

G-SHARP: Gaussian Surgical Hardware Accelerated Real-time Pipeline

Model ReleasesDGX agent

arXiv:2512.02482v2 Announce Type: replace Abstract: We propose G-SHARP, a commercially compatible, real-time surgical scene reconstruction framework designed for minimally invasive procedures that req

Gemini Live Agent Challenge: Announcing the winners and highlights

Model ReleasesDGX agent

The Gemini Live Agent Challenge is officially in the books! We challenged developers worldwide to break out of the traditional 'text box' paradigm by building next-generation AI agents. From our initi

Gemma-4-31B-it-Pearl supports 256K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powere…

Model ReleasesDGX agent

Gemma-4-31B-it-Pearl supports 256K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powered endpoint. Eventually, we plan to expand our Pearl powered

Gemma-4-31B-it-Pearl supports text and image input, 256K context, configurable thinking, function calling, and JSON mode. This is Together A…

Model ReleasesDGX agent

Gemma-4-31B-it-Pearl supports text and image input, 256K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powered endpoint. Eventually, we plan to exp

GenCircuit-RL: Reinforcement Learning from Hierarchical Verification for Genetic Circuit Design

Model ReleasesDGX agent

arXiv:2605.14215v1 Announce Type: new Abstract: Genetic circuit design remains a laborious, expert-driven process despite decades of progress in synthetic biology. We study this problem through code g

GenExam: A Multidisciplinary Text-to-Image Exam

Model ReleasesDGX agent

arXiv:2509.14232v5 Announce Type: replace Abstract: Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style bench

GFMate: Empowering Graph Foundation Models with Test-time Prompt Tuning

Model ReleasesDGX agent

arXiv:2605.14809v1 Announce Type: new Abstract: Graph prompt tuning has shown great potential in graph learning by introducing trainable prompts to enhance the model performance in conventional single

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

Model ReleasesDGX agent

arXiv:2602.06718v2 Announce Type: replace-cross Abstract: Citations provide the basis for trusting scientific claims; when they are invalid or fabricated, this trust collapses. With the advent of Larg

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

Model ReleasesDGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

Model ReleasesDGX agent

arXiv:2605.14841v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) has become the dominant paradigm for parameter-efficient fine-tuning (PEFT) of large language models (LLMs). However, its b

GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration

Model ReleasesDGX agent

arXiv:2605.13848v1 Announce Type: new Abstract: Agentic LLM frameworks that rely on prompted orchestration, where the model itself determines workflow transitions, often suffer from hallucinated routi

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Model ReleasesDGX agent

arXiv:2605.14498v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems t

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

Model ReleasesDGX agent

arXiv:2605.14821v1 Announce Type: new Abstract: Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit

Herculean: An Agentic Benchmark for Financial Intelligence

Model ReleasesDGX agent

arXiv:2605.14355v1 Announce Type: new Abstract: As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carr

Hidden State Poisoning Attacks against Mamba-based Language Models

Model ReleasesDGX agent

arXiv:2601.01972v4 Announce Type: cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. Yet, their advers

HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

Model ReleasesDGX agent

arXiv:2605.15024v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images

Holistic Evaluation and Failure Diagnosis of AI Agents

Model ReleasesDGX agent

arXiv:2605.14865v1 Announce Type: new Abstract: AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, an

How business operations teams use Codex

Model ReleasesDGX agent

This document from OpenAI describes how business operations teams leverage Codex, OpenAI's code generation model, to automate and streamline their workflows. It likely covers practical applications su

How data science teams use Codex

Model ReleasesDGX agent

This OpenAI Academy resource explores practical applications of Codex, their AI code generation model, within data science workflows and teams. It likely covers how data scientists leverage Codex to a

← Previous
1…246247248249250…377
Next →