AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

DGX agent

arXiv:2506.08136v3 Announce Type: replace Abstract: We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The ben

model-releasesarxiv-cs-cl
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

FORTIS: Benchmarking Over-Privilege in Agent Skills

DGX agent

arXiv:2605.09163v1 Announce Type: new Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This

model-releasesarxiv-cs-ai
12 May 2026
Applications

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

DGX agent

arXiv:2605.10834v1 Announce Type: new Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perfor

applicationsarxiv-cs-ai
12 May 2026
Agents

GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents

DGX agent

arXiv:2510.13896v2 Announce Type: replace-cross Abstract: Cellular image segmentation is essential for quantitative biology yet remains difficult due to heterogeneous modalities, morphological variabi

agentsarxiv-cs-ai
12 May 2026
Model Releases

MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

DGX agent

arXiv:2605.09684v1 Announce Type: cross Abstract: We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-eli

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

DGX agent

arXiv:2605.09703v1 Announce Type: new Abstract: Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on

model-releasesarxiv-cs-cv
12 May 2026
Safety

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

DGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

safetyarxiv-cs-ai
12 May 2026
Model Releases

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

DGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

DGX agent

arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W

model-releasesarxiv-cs-lg
12 May 2026
Safety

Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents

DGX agent

arXiv:2605.08964v1 Announce Type: new Abstract: In this thesis, we develop algorithms with theoretical guarantees for ensuring reliability and accountability of Machine Learning (ML) systems. As ML sy

safetyarxiv-cs-lg
12 May 2026
Local Ai

UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection

DGX agent

arXiv:2601.23273v2 Announce Type: replace Abstract: Prompt agents have recently emerged as a promising paradigm for automated prompt optimization, framing prompt discovery as a sequential decision-mak

local-aiarxiv-cs-cl
12 May 2026
Agents

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models

DGX agent

arXiv:2605.10106v1 Announce Type: cross Abstract: Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training

agentsarxiv-cs-ai
12 May 2026
Agents

ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning

DGX agent

arXiv:2605.07103v1 Announce Type: new Abstract: Reaction feasibility prediction, as a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in art

agentsarxiv-cs-ai
11 May 2026
Safety

Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents

DGX agent

arXiv:2601.21699v2 Announce Type: replace Abstract: Multi-turn reasoning agents solve complex questions by decomposing them into intermediate retrieval or tool-use steps, for accumulating supporting e

safetyarxiv-cs-cl
11 May 2026
Agents

Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response

DGX agent

arXiv:2603.02274v2 Announce Type: replace-cross Abstract: Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological r

agentsarxiv-cs-ai
11 May 2026
Agents

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

DGX agent

arXiv:2605.08083v1 Announce Type: new Abstract: Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during infe

agentsarxiv-cs-cl
11 May 2026
Agents

MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments

DGX agent

arXiv:2605.07058v1 Announce Type: cross Abstract: Real-world clinical diagnosis is a complex process in which the doctor is required to obtain information from both interaction with the patient and co

agentsarxiv-cs-ai
11 May 2026
Safety

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

DGX agent

arXiv:2604.03675v2 Announce Type: replace Abstract: Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforceme

safetyarxiv-cs-ai
11 May 2026
Agents

RelAgent: LLM Agents as Data Scientists for Relational Learning

DGX agent

arXiv:2605.07840v1 Announce Type: new Abstract: Relational learning is a challenging problem that has motivated a wide range of approaches, including graph-based models (e.g., graph neural networks, g

agentsarxiv-cs-lg
11 May 2026
Safety

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

DGX agent

arXiv:2605.06230v2 Announce Type: replace Abstract: As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool

safetyarxiv-cs-ai
11 May 2026
Model Releases

When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents

DGX agent

arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c

model-releasesarxiv-cs-cl
11 May 2026
Safety

A Skill-Based AI Agentic Pipeline for Library of Congress Subject Indexing

DGX agent

arXiv:2605.03537v1 Announce Type: cross Abstract: This paper presents a modular AI agentic skill pipeline for automating subject indexing with Library of Congress Subject Headings (LCSH). Subject inde

safetyarxiv-cs-ai
7 May 2026
Model Releases

Closed-Loop Vision-Language Planning for Multi-Agent Coordination

DGX agent

arXiv:2502.10148v3 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) struggles with sample efficiency, interpretability, and generalization. While Large Language M

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

DGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

model-releasesarxiv-cs-ai
7 May 2026
Agents

GeoDecider: A Coarse-to-Fine Agentic Workflow for Explainable Lithology Classification

DGX agent

arXiv:2605.03383v1 Announce Type: new Abstract: Lithology classification aims to infer subsurface rock types from well-logging signals, supporting downstream applications like reservoir characterizati

agentsarxiv-cs-ai
7 May 2026
Agents

Multi-Agent Strategic Games with LLMs

DGX agent

arXiv:2605.03604v1 Announce Type: cross Abstract: This paper asks whether large language models (LLMs) can be used to study the strategic foundations of conflict and cooperation. I introduce LLMs as e

agentsarxiv-cs-ai
7 May 2026
Model Releases

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

DGX agent

arXiv:2605.04019v1 Announce Type: new Abstract: AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense

DGX agent

arXiv:2605.03034v1 Announce Type: new Abstract: Agentic systems involved in high-stake decision-making under adversarial pressure need formal guarantees not offered by existing approaches. Motivated b

model-releasesarxiv-cs-ai
7 May 2026
Agents

AI Agents for Sustainable SMEs: A Green ESG Assessment Framework

DGX agent

arXiv:2605.00841v1 Announce Type: new Abstract: This study presents a novel, AI-driven framework for assessing Environmental, Social, and Governance (ESG) performance in European small and medium-size

agentsarxiv-cs-ai
6 May 2026
Safety

Architectural Obsolescence of Unhardened Agentic-AI Runtimes

DGX agent

arXiv:2605.01740v1 Announce Type: cross Abstract: An agentic-AI runtime issues tool calls, sends messages, and actuates devices on behalf of an LLM. Catching the four ways an action can diverge from i

safetyarxiv-cs-ai
6 May 2026
Agents

GDPR Auto-Formalization with AI Agents and Human Verification

DGX agent

arXiv:2604.14607v2 Announce Type: replace Abstract: We study the overall process of automatic formalization of GDPR provisions using large language models, within a human-in-the-loop verification fram

agentsarxiv-cs-ai
6 May 2026
Model Releases

LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

DGX agent

arXiv:2605.01394v1 Announce Type: cross Abstract: Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Alth

model-releasesarxiv-cs-ai
6 May 2026
Safety

MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration

DGX agent

arXiv:2605.01507v1 Announce Type: new Abstract: Prior studies report that partial driving automation can increase the cognitive demands on human drivers. This effect largely arises from human drivers'

safetyarxiv-cs-ai
6 May 2026
Agents

Practical Limits of Autonomous Test Repair: A Multi-Agent Case Study with LLM-Driven Discovery and Self-Correction

DGX agent

arXiv:2605.01471v1 Announce Type: cross Abstract: Maintaining reliable UI test suites in large-scale enterprise applications is a persistent and costly challenge. We present an industrial case study o

agentsarxiv-cs-ai
6 May 2026
Model Releases

The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate

DGX agent

arXiv:2605.00914v1 Announce Type: cross Abstract: Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

DGX agent

arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

When Alignment Isn't Enough: Response-Path Attacks on LLM Agents

DGX agent

arXiv:2605.02187v1 Announce Type: cross Abstract: Bring-Your-Own-Key (BYOK) agent architectures let users route LLM traffic through third-party relays, creating a critical integrity gap: a malicious r

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems

DGX agent

arXiv:2605.02463v1 Announce Type: cross Abstract: Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, t

model-releasesarxiv-cs-ai
6 May 2026
Agents

Autonomous QA Agent: A Retrieval-Augmented Framework for Reliable Selenium Script Generation

DGX agent

arXiv:2601.06034v1 Announce Type: cross Abstract: Software testing is critical in the software development lifecycle, yet translating requirements into executable test scripts remains manual and error

agentsarxiv-cs-lg
5 May 2026
Model Releases

CP-SynC: Multi-Agent Zero-Shot Constraint Modeling in MiniZinc with Synthesized Checkers

DGX agent

arXiv:2605.01675v1 Announce Type: cross Abstract: Constraint Programming (CP) is a powerful paradigm for solving combinatorial problems, yet translating natural language problem descriptions into exec

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture

DGX agent

arXiv:2605.01567v1 Announce Type: cross Abstract: Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineeri

model-releasesarxiv-cs-cl
5 May 2026
Agents

GA-VisAgent: A Multi-Agent application for code generation and visualization in interactive learning

DGX agent

arXiv:2605.01299v1 Announce Type: new Abstract: Geometric Algebra (GA) presents challenges to learners due to its highly abstract mathematical structure and complex operational rules, as translating a

agentsarxiv-cs-lg
5 May 2026
Model Releases

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

DGX agent

arXiv:2505.10887v3 Announce Type: replace Abstract: This paper introduces extsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, i

model-releasesarxiv-cs-ai
5 May 2026
Agents

LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery

DGX agent

arXiv:2508.12232v3 Announce Type: replace-cross Abstract: Issue-to-commit link recovery in software repositories is fundamental to software traceability and project management, yet it remains a challe

agentsarxiv-cs-ai
5 May 2026
Model Releases

MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding

DGX agent

arXiv:2510.08804v3 Announce Type: replace Abstract: We present MOSAIC, a multi-agent Large Language Model (LLM) framework for solving challenging scientific coding tasks. Unlike general-purpose coding

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Planner Matters! An Efficient and Unbalanced Multi-agent Collaboration Framework for Long-horizon Planning

DGX agent

arXiv:2605.02168v1 Announce Type: cross Abstract: Language model (LM)-based agents have demonstrated promising capabilities in automating complex tasks from natural language instructions, yet they con

model-releasesarxiv-cs-lg
5 May 2026
Safety

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

DGX agent

arXiv:2605.01416v1 Announce Type: cross Abstract: The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy

safetyarxiv-cs-cl
5 May 2026
Agents

AgentEconomist: An End-to-end Agentic System Translating Economic Intuitions into Executable Computational Experiments

DGX agent

arXiv:2604.27725v1 Announce Type: cross Abstract: A long-standing challenge in economics lies not in the lack of intuition, but in the difficulty of translating intuitive insights into verifiable rese

agentsarxiv-cs-ai
1 May 2026
← Previous
1…8283848586…236
Next →