AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Safety

CARL: Criticality-Aware Agentic Reinforcement Learning

DGX agent

arXiv:2512.04949v3 Announce Type: replace-cross Abstract: Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction.

safetyarxiv-cs-ai
12 May 2026
Agents

CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.08399v1 Announce Type: new Abstract: Tool-augmented language models can extend small language models with external executable skills, but scaling the tool library creates a coupled challeng

agentsarxiv-cs-ai
12 May 2026
Model Releases

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

DGX agent

arXiv:2605.09675v1 Announce Type: new Abstract: Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%

DGX agent

arXiv:2605.08112v1 Announce Type: cross Abstract: AI coding agents powered by large language models can read codebases and produce functional code, but they routinely violate team-specific product dec

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

DGX agent

arXiv:2605.08747v1 Announce Type: new Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call te

model-releasesarxiv-cs-ai
12 May 2026
Safety

EGL-SCA: Structural Credit Assignment for Co-Evolving Instructions and Tools in Graph Reasoning Agents

DGX agent

arXiv:2605.10366v1 Announce Type: new Abstract: Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text,

safetyarxiv-cs-ai
12 May 2026
Model Releases

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

DGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

model-releasesarxiv-cs-cl
12 May 2026
Safety

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL

DGX agent

arXiv:2511.01008v2 Announce Type: replace Abstract: Large Language Models (LLMs) often struggle with the precise logic and schema alignment required for complex Text-to-SQL tasks. While current method

safetyarxiv-cs-cl
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

DGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

model-releasesarxiv-cs-cl
12 May 2026
Research

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

DGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

researcharxiv-cs-ai
12 May 2026
Model Releases

Priority-Driven Control and Communication in Decentralized Multi-Agent Systems via Reinforcement Learning

DGX agent

arXiv:2605.10482v1 Announce Type: cross Abstract: Event-triggered control provides a mechanism for avoiding excessive use of constrained communication bandwidth in networked multi-agent systems. Howev

model-releasesarxiv-cs-lg
12 May 2026
Safety

SceneFactory: GPU-Accelerated Multi-Agent Driving Simulation with Physics-Based Vehicle Dynamics

DGX agent

arXiv:2605.08528v1 Announce Type: cross Abstract: Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide

safetyarxiv-cs-ro
12 May 2026
Research

The Invisible Handshake: Persistent Overpricing by Adaptive Market Agents

DGX agent

arXiv:2510.15995v3 Announce Type: replace-cross Abstract: We study overpricing in a repeated game between two representative agents: a market maker, who controls market liquidity, and a market taker,

researcharxiv-cs-lg
12 May 2026
Agents

TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning

DGX agent

arXiv:2605.10038v1 Announce Type: new Abstract: Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires

agentsarxiv-cs-ai
12 May 2026
Safety

TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System

DGX agent

arXiv:2602.03688v2 Announce Type: replace Abstract: Multi-round LLM-based multi-agent systems rely on effective communication structures to support collaboration across rounds. However, most existing

safetyarxiv-cs-ai
12 May 2026
Model Releases

Unpredictability dissociates from structured control in language agents

DGX agent

arXiv:2605.09692v1 Announce Type: new Abstract: Unpredictable behavior is often taken as evidence of control, yet stochastic dispersion and structured action control need not coincide. This paper test

model-releasesarxiv-cs-ai
12 May 2026
Local Ai

Verifiable Process Rewards for Agentic Reasoning

DGX agent

arXiv:2605.10325v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches

local-aiarxiv-cs-ai
12 May 2026
Model Releases

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents

DGX agent

arXiv:2605.06607v2 Announce Type: replace-cross Abstract: Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemi

model-releasesarxiv-cs-ai
11 May 2026
Agents

FlightSense: An End-to-End MLOps Platform for Real-Time Flight Delay Prediction via Rotation-Chain Propagation Features and Agentic Conversational AI

DGX agent

arXiv:2605.07364v1 Announce Type: new Abstract: Flight delays impose cascading operational and financial burdens across the aviation network, costing the U.S. economy billions of dollars annually by d

agentsarxiv-cs-lg
11 May 2026
Local Ai

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization

DGX agent

arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp

local-aiarxiv-cs-ai
11 May 2026
Model Releases

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

DGX agent

arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H

model-releasesarxiv-cs-ai
11 May 2026
Local Ai

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

DGX agent

arXiv:2605.07276v1 Announce Type: new Abstract: Code-agent RL often receives weak feedback: rollout-time signals are reliable and executable, but capture only necessary or surface conditions for task

local-aiarxiv-cs-ai
11 May 2026
Safety

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

DGX agent

arXiv:2605.06130v2 Announce Type: replace Abstract: A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupl

safetyarxiv-cs-ai
11 May 2026
Model Releases

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

DGX agent

arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

DGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Agentic Vulnerability Reasoning on Windows COM Binaries

DGX agent

arXiv:2605.05000v1 Announce Type: cross Abstract: Windows Component Object Model (COM) services run with elevated privileges and are widely accessible to authenticated users, making race conditions in

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

DGX agent

arXiv:2511.02230v4 Announce Type: replace-cross Abstract: KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV

model-releasesarxiv-cs-ai
7 May 2026
Local Ai

FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing

DGX agent

arXiv:2506.20911v2 Announce Type: replace Abstract: We develop a cost-efficient neurosymbolic agent to address challenging multi-turn image editing tasks such as ``Detect the bench in the image while

local-aiarxiv-cs-cv
7 May 2026
Safety

Graph-SND: Sparse Aggregation for Behavioral Diversity in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.05020v1 Announce Type: new Abstract: System Neural Diversity (SND) measures behavioral heterogeneity in multi-agent reinforcement learning by averaging pairwise distances over all inom{n}{2

safetyarxiv-cs-lg
7 May 2026
Model Releases

MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents

DGX agent

arXiv:2605.03675v1 Announce Type: new Abstract: Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

QKVShare: Quantized KV-Cache Handoff for Multi-Agent On-Device LLMs

DGX agent

arXiv:2605.03884v1 Announce Type: new Abstract: Multi-agent LLM systems on edge devices need to hand off latent context efficiently, but the practical choices today are expensive re-prefill or full-pr

model-releasesarxiv-cs-ai
7 May 2026
Local Ai

ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting

DGX agent

arXiv:2605.03804v1 Announce Type: new Abstract: Long-term personalized memory for LLM agents is challenging on resource-limited edge devices due to high storage costs and multimodal complexity. To add

local-aiarxiv-cs-ai
7 May 2026
Model Releases

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

DGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

model-releasesarxiv-cs-cl
7 May 2026
Safety

ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms

DGX agent

arXiv:2605.03212v1 Announce Type: cross Abstract: Modeling latent clinical constructs from unconstrained clinical interactions is a unique challenge in affective computing. We present ADAPTS (Agentic

safetyarxiv-cs-cl
6 May 2026
Model Releases

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework

DGX agent

arXiv:2605.01604v1 Announce Type: new Abstract: Existing evaluation frameworks for large language models -- including HELM, MT-Bench, AgentBench, and BIG-bench -- are designed for controlled, single-s

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level

DGX agent

arXiv:2601.03731v3 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve into autonomous agents, evaluating repository-level reasoning, the ability to maintain logical consiste

model-releasesarxiv-cs-ai
6 May 2026
Agents

From Synthesis to Clinical Assistance: A Strategy-Aware Agent Framework for Autism Intervention based on Real Clinical Dataset

DGX agent

arXiv:2605.02916v1 Announce Type: new Abstract: The development of AI-assisted Early Intensive Behavioral Intervention (EIBI) for Autism Spectrum Disorder (ASD) is severely constrained by data scarcit

agentsarxiv-cs-lg
6 May 2026
Local Ai

GISclaw: A Comprehensive Open-Source LLM Agent System for Realistic Multi-Step Geospatial Analysis

DGX agent

arXiv:2603.26845v2 Announce Type: replace-cross Abstract: Most LLM-driven GIS assistants solve narrow single-step tasks tightly coupled to proprietary platforms such as ArcGIS or QGIS, limiting their

local-aiarxiv-cs-ai
6 May 2026
Model Releases

GOAT: A Training Framework for Goal-Oriented Agent with Tools

DGX agent

arXiv:2510.12218v2 Announce Type: replace Abstract: Current approaches rely on zero-shot evaluation due to the absence of training data; while proprietary models such as GPT-4 exhibit strong reasoning

model-releasesarxiv-cs-ai
6 May 2026
Research

MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents

DGX agent

arXiv:2605.03482v1 Announce Type: cross Abstract: Persistent external memory enables LLM agents to maintain context across sessions, yet its security properties remain formally uncharacterized. We for

researcharxiv-cs-lg
6 May 2026
Model Releases

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

DGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

DGX agent

arXiv:2605.00425v1 Announce Type: new Abstract: Reinforcement learning (RL) has significantly advanced the ability of large language model (LLM) agents to interact with environments and solve multi-tu

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling

DGX agent

arXiv:2605.00833v1 Announce Type: new Abstract: Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Exis

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Automated Interpretability and Feature Discovery in Language Models with Agents

DGX agent

arXiv:2605.01555v1 Announce Type: new Abstract: We introduce an autonomous multiagent framework for mechanistic interpretability that automates both explaining and finding internal features in large l

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Beating the Style Detector: Three Hours of Agentic Research on the AI-Text Arms Race

DGX agent

arXiv:2605.02620v1 Announce Type: new Abstract: Reproducing an empirical NLP study used to take weeks. Given the released data and a modern agentic-research harness, we redo every experiment of a rece

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

DGX agent

arXiv:2605.01371v1 Announce Type: new Abstract: The rapid advancement of Multimodal Large Language Models (MLLMs) has empowered Unmanned Aerial Vehicle (UAV) with exceptional capabilities in spatial r

model-releasesarxiv-cs-ro
5 May 2026
Model Releases

FeedbackLLM: Metadata driven Multi-Agentic Language Agnostic Test Case Generator with Evolving prompt and Coverage Feedback

DGX agent

arXiv:2605.01264v1 Announce Type: cross Abstract: Traditional approaches to test case generation often involve manual effort and incur significant computational overhead. Additionally, these approache

model-releasesarxiv-cs-lg
5 May 2026
← Previous
1…9293949596…236
Next →