AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Safety

From surveillance to signalling: escalation channels as environmental controls for agentic AI

DGX agent

arXiv:2510.05192v2 Announce Type: replace-cross Abstract: When AI agents operating with access to sensitive information encounter a conflict between completing an assigned task and following rules or

safetyarxiv-cs-ai
1 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

DGX agent

arXiv:2604.27934v1 Announce Type: new Abstract: Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting sign

agentsarxiv-cs-ai
1 May 2026
Agents

Think it, Run it: Autonomous ML pipeline generation via self-healing multi-agent AI

DGX agent

arXiv:2604.27096v1 Announce Type: new Abstract: The purpose of our paper is to develop a unified multi-agent architecture that automates end-to-end machine learning (ML) pipeline generation from datas

agentsarxiv-cs-ai
1 May 2026
Agents

Trace-Level Analysis of Information Contamination in Multi-Agent Systems

DGX agent

arXiv:2604.27586v1 Announce Type: new Abstract: Reasoning over heterogeneous artifacts (PDFs, spreadsheets, slide decks, etc.) increasingly occurs within structured agent workflows that iteratively ex

agentsarxiv-cs-ai
1 May 2026
Agents

EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks

DGX agent

arXiv:2502.05907v3 Announce Type: replace Abstract: Completing Long-Horizon (LH) tasks in open-ended worlds is an important yet difficult problem for embodied agents. Existing approaches suffer from t

agentsarxiv-cs-ro
30 Apr 2026
Agents

OxyGent: Making Multi-Agent Systems Modular, Observable, and Evolvable via Oxy Abstraction

DGX agent

arXiv:2604.25602v2 Announce Type: replace Abstract: Deploying production-ready multi-agent systems (MAS) in complex industrial environments remains challenging due to limitations in scalability, obser

agentsarxiv-cs-ai
30 Apr 2026
Agents

Training Computer Use Agents to Assess the Usability of Graphical User Interfaces

DGX agent

arXiv:2604.26020v1 Announce Type: cross Abstract: Usability testing with experts and potential users can assess the effectiveness, efficiency, and user satisfaction of graphical user interfaces (GUIs)

agentsarxiv-cs-ai
30 Apr 2026
Model Releases

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

DGX agent

arXiv:2604.24955v1 Announce Type: new Abstract: As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken

model-releasesarxiv-cs-cl
29 Apr 2026
Safety

Failure-Centered Runtime Evaluation for Deployed Trilingual Public-Space Agents

DGX agent

arXiv:2604.23990v1 Announce Type: new Abstract: This paper presents PSA-Eval, a failure-centered runtime evaluation framework for deployed trilingual public-space agents. The central claim is that, wh

safetyarxiv-cs-ai
28 Apr 2026
Tutorials

From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents

DGX agent

arXiv:2604.23194v1 Announce Type: new Abstract: Large language model-based agents have recently emerged as powerful approaches for solving dynamic and multi-step tasks. Most existing agents employ pla

tutorialsarxiv-cs-ai
28 Apr 2026
Model Releases

Latency and Cost of Multi-Agent Intelligent Tutoring at Scale

DGX agent

arXiv:2604.24110v1 Announce Type: cross Abstract: Multi-agent LLM tutoring systems improve response quality through agent specialization, but each student query triggers several concurrent API calls w

model-releasesarxiv-cs-ai
28 Apr 2026
Safety

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

DGX agent

arXiv:2604.24348v1 Announce Type: new Abstract: The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS

safetyarxiv-cs-cl
28 Apr 2026
Model Releases

Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably

DGX agent

arXiv:2603.18563v2 Announce Type: replace Abstract: As autonomous AI agents increasingly mediate online platform markets, a fundamental question emerges: do these markets generate stable strategic out

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems

DGX agent

arXiv:2508.10177v3 Announce Type: replace Abstract: Recent Large Language Model (LLM)-based AutoML systems demonstrate impressive capabilities but face significant limitations such as constrained expl

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations

DGX agent

arXiv:2604.20848v1 Announce Type: cross Abstract: Large Language Model (LLM)-based recommendation systems have demonstrated remarkable capabilities in understanding user preferences and generating per

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models

DGX agent

arXiv:2604.21896v1 Announce Type: new Abstract: This paper introduces a new paradigm for AI game programming, leveraging large language models (LLMs) to extend and operationalize Claude Shannon's taxo

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

DGX agent

arXiv:2604.20200v1 Announce Type: new Abstract: Frontier coding agents are increasingly used in workflows where users supervise progress primarily through repeated improvement of a public score, namel

model-releasesarxiv-cs-cl
23 Apr 2026
Agents

Enhancing Research Idea Generation through Combinatorial Innovation and Multi-Agent Iterative Search Strategies

DGX agent

arXiv:2604.20548v1 Announce Type: cross Abstract: Scientific progress depends on the continual generation of innovative re-search ideas. However, the rapid growth of scientific literature has greatly

agentsarxiv-cs-ai
23 Apr 2026
Model Releases

From Data to Theory: Autonomous Large Language Model Agents for Materials Science

DGX agent

arXiv:2604.19789v1 Announce Type: new Abstract: We present an autonomous large language model (LLM) agent for end-to-end, data-driven materials theory development. The model can choose an equation for

model-releasesarxiv-cs-ai
23 Apr 2026
Agents

Mol-Debate: Multi-Agent Debate Improves Structural Reasoning in Molecular Design

DGX agent

arXiv:2604.20254v1 Announce Type: new Abstract: Text-guided molecular design is a key capability for AI-driven drug discovery, yet it remains challenging to map sequential natural-language instruction

agentsarxiv-cs-ai
23 Apr 2026
Hardware

ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants

DGX agent

arXiv:2604.18616v1 Announce Type: cross Abstract: LLM-based coding agents can generate functionally correct GPU kernels, yet their performance remains far below hand-optimized libraries on critical co

hardwarearxiv-cs-ai
22 Apr 2026
Safety

BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search

DGX agent

arXiv:2601.11037v2 Announce Type: replace Abstract: RL-based agentic search enables LLMs to solve complex questions via dynamic planning and external search. While this approach significantly enhances

safetyarxiv-cs-ai
22 Apr 2026
Model Releases

Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?

DGX agent

arXiv:2602.18571v2 Announce Type: replace-cross Abstract: While significant progress has been made in automating various aspects of software development through coding agents, there is still significa

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents

DGX agent

arXiv:2604.19457v1 Announce Type: new Abstract: Long-horizon enterprise agents make high-stakes decisions (loan underwriting, claims adjudication, clinical review, prior authorization) under lossy mem

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

How Adversarial Environments Mislead Agentic AI?

DGX agent

arXiv:2604.18874v1 Announce Type: new Abstract: Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Human-Guided Harm Recovery for Computer Use Agents

DGX agent

arXiv:2604.18847v1 Announce Type: new Abstract: As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectivel

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training

DGX agent

arXiv:2510.12831v3 Announce Type: replace Abstract: Multi-turn Text-to-SQL aims to translate a user's conversational utterances into executable SQL while preserving dialogue coherence and grounding to

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

DGX agent

arXiv:2604.18557v1 Announce Type: new Abstract: Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexi

safetyarxiv-cs-cv
21 Apr 2026
Safety

AI Agents and Hard Choices

DGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

safetyarxiv-cs-ai
20 Apr 2026
Model Releases

COMPASS: Benchmarking Constrained Optimization in LLM Agents

DGX agent

arXiv:2510.07043v2 Announce Type: replace Abstract: Human decision-making often involves constrained optimization. As LLM agents are deployed to assist with real-world tasks like travel planning, shop

model-releasesarxiv-cs-lg
20 Apr 2026
Safety

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents

DGX agent

arXiv:2504.13541v5 Announce Type: replace-cross Abstract: Training resource-constrained autonomous agents on multiple tasks simultaneously is crucial for adapting to diverse real-world environments. R

safetyarxiv-cs-ai
20 Apr 2026
Local Ai

Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap

DGX agent

arXiv:2604.15075v1 Announce Type: cross Abstract: Open-weight Small Language Models(SLMs) can provide faster local inference at lower financial cost, but may not achieve the same performance level as

local-aiarxiv-cs-lg
17 Apr 2026
Safety

Exploration and Exploitation Errors Are Measurable for Language Model Agents

DGX agent

arXiv:2604.13151v1 Announce Type: new Abstract: Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these

safetyarxiv-cs-ai
17 Apr 2026
Safety

ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

DGX agent

arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg

safetyarxiv-cs-ai
17 Apr 2026
Agents

Towards a Multi-Embodied Grasping Agent

DGX agent

arXiv:2510.27420v3 Announce Type: replace Abstract: Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often l

agentsarxiv-cs-ro
17 Apr 2026
Model Releases

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

DGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

model-releasesarxiv-cs-lg
16 Apr 2026
Model Releases

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

DGX agent

arXiv:2512.20798v4 Announce Type: replace Abstract: As autonomous AI agents are deployed in high-stakes environments, ensuring their safety has become a paramount concern. Existing safety benchmarks p

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

DGX agent

arXiv:2604.12762v1 Announce Type: cross Abstract: We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an ag

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

From Plan to Action: How Well Do Agents Follow the Plan?

DGX agent

arXiv:2604.12147v1 Announce Type: cross Abstract: Agents aspire to eliminate the need for task-specific prompt crafting through autonomous reason-act-observe loops. Still, they are commonly instructed

model-releasesarxiv-cs-ai
15 Apr 2026
Agents

Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning

DGX agent

arXiv:2604.12282v1 Announce Type: new Abstract: Spreadsheets are central to real-world applications such as enterprise reporting, auditing, and scientific data management. Despite their ubiquity, exis

agentsarxiv-cs-cl
15 Apr 2026
Model Releases

Competing with AI Scientists: Agent-Driven Approach to Astrophysics Research

DGX agent

arXiv:2604.09621v1 Announce Type: new Abstract: We present an agent-driven approach to the construction of parameter inference pipelines for scientific data analysis. Our method leverages a multi-agen

model-releasesarxiv-cs-ai
14 Apr 2026
Agents

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning

DGX agent

arXiv:2604.09813v1 Announce Type: new Abstract: Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable envir

agentsarxiv-cs-ai
14 Apr 2026
Agents

DarwinNet: An Evolutionary Network Architecture for Agent-Driven Protocol Synthesis

DGX agent

arXiv:2604.01236v2 Announce Type: replace-cross Abstract: Traditional network architectures suffer from severe protocol ossification and structural fragility due to their reliance on static, human-def

agentsarxiv-cs-ai
14 Apr 2026
Model Releases

Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems

DGX agent

arXiv:2604.09666v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) and its graph-based extensions (GraphRAG) are effective paradigms for improving large language model (LLM) reason

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution

DGX agent

arXiv:2604.09568v1 Announce Type: cross Abstract: High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant chall

model-releasesarxiv-cs-cl
14 Apr 2026
Local Ai

MAFIG: Multi-agent Driven Formal Instruction Generation Framework

DGX agent

arXiv:2604.10989v1 Announce Type: new Abstract: Emergency situations in scheduling systems often trigger local functional failures that undermine system stability and even cause system collapse. Exist

local-aiarxiv-cs-ai
14 Apr 2026
Model Releases

The Amazing Agent Race: Strong Tool Users, Weak Navigators

DGX agent

arXiv:2604.10261v1 Announce Type: new Abstract: Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents

DGX agent

arXiv:2604.11465v1 Announce Type: new Abstract: Large language model (LLM) agents show promise on realistic tool-use tasks, but deploying capable agents on modest hardware remains challenging. We stud

model-releasesarxiv-cs-ai
14 Apr 2026
← Previous
1…5253545556…233
Next →