AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Local Ai

SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery

DGX agent

arXiv:2607.02807v1 Announce Type: new Abstract: Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems. However, they tend to converge onto a s

local-aiarxiv-cs-ai
7 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Toward Efficient Agents: Memory, Tool learning, and Planning

DGX agent

arXiv:2601.14192v2 Announce Type: replace Abstract: Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has conti

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

DGX agent

arXiv:2607.04425v1 Announce Type: cross Abstract: Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform int

safetyarxiv-cs-ai
7 Jul 2026
Safety

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

DGX agent

arXiv:2607.05132v1 Announce Type: cross Abstract: As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents tha

safetyarxiv-cs-cl
7 Jul 2026
Agents

A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development

DGX agent

arXiv:2603.04390v2 Announce Type: replace Abstract: WebGIS development requires consistency, yet agentic AI often fails due to LLM context constraints, forgetting, stochasticity, instruction failure,

agentsarxiv-cs-ai
3 Jul 2026
Model Releases

BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems

DGX agent

arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_

model-releasesarxiv-cs-cl
3 Jul 2026
Agents

Coding-agents can replicate scientific machine learning papers

DGX agent

arXiv:2607.02134v1 Announce Type: new Abstract: Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predic

agentsarxiv-cs-ai
3 Jul 2026
Agents

MMAO-Cls: Metabolic Multi-Agent Optimization for Joint Feature Selection and Classifier Tuning

DGX agent

arXiv:2607.01539v1 Announce Type: cross Abstract: This paper studies whether the Metabolic Multi-Agent Optimizer (MMAO) can act as a credible outer-loop optimizer for classification model selection. W

agentsarxiv-cs-lg
3 Jul 2026
Agents

Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection

DGX agent

arXiv:2607.01367v1 Announce Type: cross Abstract: A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MARL). While MARL has already been employe

agentsarxiv-cs-ro
3 Jul 2026
Model Releases

Steerability via constraints: a substrate for scalable oversight of coding agents

DGX agent

arXiv:2607.02389v1 Announce Type: new Abstract: Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human

model-releasesarxiv-cs-ai
3 Jul 2026
Agents

The Rollout Infrastructure Tax in Coding-Agent Reinforcement Learning

DGX agent

arXiv:2607.01415v1 Announce Type: new Abstract: Coding-agent reinforcement learning treats execution infrastructure as a background implementation detail, despite relying on large numbers of interacti

agentsarxiv-cs-lg
3 Jul 2026
Agents

BaRA: BFS-and-Reflection Web Data Collection Agent

DGX agent

arXiv:2607.00007v1 Announce Type: cross Abstract: Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, ret

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

GameDevBench: Evaluating Agentic Capabilities Through Game Development

DGX agent

arXiv:2602.11103v2 Announce Type: replace Abstract: Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation

model-releasesarxiv-cs-ai
2 Jul 2026
Agents

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

DGX agent

arXiv:2601.04424v2 Announce Type: replace Abstract: Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclea

agentsarxiv-cs-cl
2 Jul 2026
Agents

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection

DGX agent

arXiv:2607.00035v1 Announce Type: new Abstract: LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, bro

agentsarxiv-cs-ai
2 Jul 2026
Agents

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

DGX agent

arXiv:2607.00597v1 Announce Type: new Abstract: Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, an

agentsarxiv-cs-cl
2 Jul 2026
Agents

NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents

DGX agent

arXiv:2601.14660v2 Announce Type: replace-cross Abstract: Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enabling t

agentsarxiv-cs-ai
2 Jul 2026
Agents

Self-GC: Self-Governing Context for Long-Horizon LLM Agents

DGX agent

arXiv:2607.00692v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. C

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests

DGX agent

arXiv:2607.00990v1 Announce Type: cross Abstract: Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue re

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Agents

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

DGX agent

arXiv:2606.19501v2 Announce Type: replace Abstract: Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read

agentsarxiv-cs-ai
1 Jul 2026
Agents

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

DGX agent

arXiv:2606.31980v1 Announce Type: new Abstract: Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a mul

agentsarxiv-cs-cl
1 Jul 2026
Model Releases

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

DGX agent

arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant atten

model-releasesarxiv-cs-ai
1 Jul 2026
Agents

OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents

DGX agent

arXiv:2606.31046v1 Announce Type: new Abstract: Artificial life has explored life-like behavior on many computational substrates, but mostly in researcher-designed closed worlds. We argue that large l

agentsarxiv-cs-ai
1 Jul 2026
Model Releases

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

DGX agent

arXiv:2606.30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on

model-releasesarxiv-cs-ai
1 Jul 2026
Agents

Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

DGX agent

arXiv:2606.30801v1 Announce Type: new Abstract: Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors

agentsarxiv-cs-cl
1 Jul 2026
Model Releases

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

DGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

DGX agent

arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Hierarchical Experimentalist Agents

DGX agent

arXiv:2606.29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametr

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning

DGX agent

arXiv:2606.28349v1 Announce Type: cross Abstract: Long-context reasoning requires models to access, retrieve, and integrate evidence scattered across documents, dialogues, and accumulated interaction

agentsarxiv-cs-ai
30 Jun 2026
Local Ai

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning

DGX agent

arXiv:2606.29425v1 Announce Type: new Abstract: Existing multi-agent debate frameworks suffer from two critical limitations: they rely on static architectures where agent roles and coordination patter

local-aiarxiv-cs-ai
30 Jun 2026
Agents

Monte Carlo Query Search: Active Capability Assessment of AI Agents

DGX agent

arXiv:2512.16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires metho

agentsarxiv-cs-ai
30 Jun 2026
Model Releases

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

DGX agent

arXiv:2603.29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visuali

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

DGX agent

arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate t

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

DGX agent

arXiv:2606.13544v3 Announce Type: replace-cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor compe

agentsarxiv-cs-ai
29 Jun 2026
Agents

Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

DGX agent

arXiv:2606.27936v1 Announce Type: cross Abstract: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by t

agentsarxiv-cs-ai
29 Jun 2026
Safety

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

DGX agent

arXiv:2510.16492v4 Announce Type: replace Abstract: As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While

safetyarxiv-cs-cl
29 Jun 2026
Agents

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

DGX agent

arXiv:2606.27974v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fi

agentsarxiv-cs-ai
29 Jun 2026
Safety

Training Observable Control Policies to Expose Agent State Through Actions

DGX agent

arXiv:2606.27609v1 Announce Type: new Abstract: Physical or operational constraints often impose communications limitations on autonomous agents. Such limitations complicate monitoring or multiagent c

safetyarxiv-cs-lg
29 Jun 2026
Agents

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

DGX agent

arXiv:2606.27251v1 Announce Type: cross Abstract: Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT)

agentsarxiv-cs-ai
26 Jun 2026
Safety

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

DGX agent

arXiv:2606.26122v1 Announce Type: new Abstract: Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples

safetyarxiv-cs-cv
26 Jun 2026
Model Releases

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

DGX agent

arXiv:2606.26356v1 Announce Type: new Abstract: Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despi

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

DGX agent

arXiv:2606.26793v1 Announce Type: cross Abstract: Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image inj

agentsarxiv-cs-ai
26 Jun 2026
Agents

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

DGX agent

arXiv:2606.26669v1 Announce Type: new Abstract: Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explo

agentsarxiv-cs-ai
26 Jun 2026
Local Ai

Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory

DGX agent

arXiv:2606.25115v1 Announce Type: new Abstract: On-device language-model agents improve by accumulating experience in retrieved memory rather than by updating weights. This memory is hard-bounded and

local-aiarxiv-cs-lg
25 Jun 2026
Agents

PhoneBuddy: Training Open Models for Agentic Phone Use

DGX agent

arXiv:2606.23049v2 Announce Type: replace Abstract: Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult bec

agentsarxiv-cs-cl
25 Jun 2026
Safety

Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See

DGX agent

arXiv:2606.25127v1 Announce Type: new Abstract: We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Per

safetyarxiv-cs-lg
25 Jun 2026
Safety

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

DGX agent

arXiv:2606.20615v2 Announce Type: replace Abstract: AI agents now participate as first-class team members across the software development lifecycle, yet no specification language exists for expressing

safetyarxiv-cs-ai
25 Jun 2026
← Previous
1…4647484950…233
Next →