AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Agents

Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents

DGX agent

arXiv:2606.10315v1 Announce Type: cross Abstract: LLM-as-judge is the default instrument for evaluating conversational agents, yet its reliability is almost always reported as agreement with human rat

agentsarxiv-cs-ai
10 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents

DGX agent

arXiv:2606.09863v1 Announce Type: new Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise. We study this failure mode, false success, across

agentsarxiv-cs-lg
10 Jun 2026
Agents

Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory

DGX agent

arXiv:2606.10677v1 Announce Type: new Abstract: Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions. Existing memory systems often s

agentsarxiv-cs-ai
10 Jun 2026
Model Releases

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

DGX agent

arXiv:2606.11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely

model-releasesarxiv-cs-ai
10 Jun 2026
Agents

ConMem: Structured Memory-Guided Adaptation in Training-Free Multi-Agent Systems

DGX agent

arXiv:2606.08702v1 Announce Type: new Abstract: Recent advances have improved the adaptive capabilities of LLM-based multi-agent systems (MAS) through memory-, skill-, and learning-based approaches, y

agentsarxiv-cs-ai
9 Jun 2026
Agents

Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents

DGX agent

arXiv:2606.07904v1 Announce Type: new Abstract: Tool-augmented large language model agents increasingly rely on external APIs, but standard tool schemas describe how to call a tool, not when the tool

agentsarxiv-cs-ai
9 Jun 2026
Model Releases

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy

DGX agent

arXiv:2606.08367v1 Announce Type: cross Abstract: Most evaluations of LLM agents look like exams: a discrete task, a clean environment, a score in minutes or hours. We argue that this approach is mism

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

Goal-Oriented Reasoning for RAG-based Memory in Conversational Agentic LLM Systems

DGX agent

arXiv:2605.12213v2 Announce Type: replace Abstract: LLM-based conversational AI agents struggle to maintain coherent behavior over long horizons due to limited context. While RAG-based approaches are

agentsarxiv-cs-ai
9 Jun 2026
Agents

SIGA: Self-Evolving Coding-Agent Adapters for Scientific Simulation

DGX agent

arXiv:2606.09774v1 Announce Type: new Abstract: Advanced scientific simulators expose specialized input languages that turn simulation goals into executable configurations, but learning them can cost

agentsarxiv-cs-ai
9 Jun 2026
Agents

Traxia: A Framework for Verifiable, Agent-Native Scientific Publishing

DGX agent

arXiv:2606.08256v1 Announce Type: new Abstract: Verifiability, attribution, and reproducibility are foundational requirements of scientific knowledge, yet current publishing infrastructure does not en

agentsarxiv-cs-ai
9 Jun 2026
Model Releases

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

DGX agent

arXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera

model-releasesarxiv-cs-cv
9 Jun 2026
Agents

Autonomous computational catalysis through an agentic research system

DGX agent

arXiv:2601.13508v4 Announce Type: replace-cross Abstract: Autonomous agents are beginning to transform scientific research from tool-assisted workflows toward self-sustaining discovery processes. Comp

agentsarxiv-cs-ai
8 Jun 2026
Agents

Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning

DGX agent

arXiv:2606.06976v1 Announce Type: new Abstract: Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct response

agentsarxiv-cs-ai
8 Jun 2026
Agents

How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope

DGX agent

arXiv:2606.07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute t

agentsarxiv-cs-ai
8 Jun 2026
Local Ai

Off-Policy Evaluation with Strategic Agents via Local Disclosure

DGX agent

arXiv:2606.07308v1 Announce Type: new Abstract: We study off-policy evaluation (OPE) under strategic behavior where decision subjects (or agents) respond to a decision maker's policy by strategically

local-aiarxiv-cs-ai
8 Jun 2026
Agents

OpenSkill: Open-World Self-Evolution for LLM Agents

DGX agent

arXiv:2606.06741v1 Announce Type: new Abstract: Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful tra

agentsarxiv-cs-ai
8 Jun 2026
Agents

Unsupervised Skill Discovery for Agentic Data Analysis

DGX agent

arXiv:2606.06416v1 Announce Type: cross Abstract: Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updati

agentsarxiv-cs-cl
5 Jun 2026
Hardware

AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning

DGX agent

arXiv:2606.04484v1 Announce Type: new Abstract: We present AgentJet, a distributed swarm training framework for large language model (LLM) agent reinforcement learning. Unlike centralized frameworks t

hardwarearxiv-cs-ai
4 Jun 2026
Agents

From Prompt to Process: a Process Taxonomy and Comparative Assessment of Frameworks Supporting AI Software Development Agents

DGX agent

arXiv:2606.04967v1 Announce Type: cross Abstract: AI tools for programming are no longer just autocomplete or chat assistants: they organize themselves as development frameworks, with process, roles,

agentsarxiv-cs-ai
4 Jun 2026
Safety

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

DGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Streaming Communication in Multi-Agent Reasoning

DGX agent

arXiv:2606.05158v1 Announce Type: cross Abstract: Multi-agent reasoning systems adopt a 'generate-then-transfer' paradigm that forces end-to-end latency to scale linearly with pipeline depth. We intro

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

DGX agent

arXiv:2602.12430v4 Announce Type: replace-cross Abstract: The transition from monolithic language models to modular, skill-equipped agents marks a defining shift in how large language models (LLMs) ar

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Co-evolving Agent Architectures and Interpretable Reasoning for Automated Optimization

DGX agent

arXiv:2604.17708v2 Announce Type: replace Abstract: Automating operations research (OR) with large language models (LLMs) remains limited by hand-crafted reasoning--execution workflows. Complex OR tas

agentsarxiv-cs-ai
3 Jun 2026
Agents

FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration

DGX agent

arXiv:2512.11213v2 Announce Type: replace Abstract: Scaling test-time computation has been shown to significantly improve large language model (LLM) performance without additional training. However, e

agentsarxiv-cs-ai
3 Jun 2026
Agents

Uncertainty-Aware Clarification in LLM Agents with Information Gain

DGX agent

arXiv:2606.03135v1 Announce Type: new Abstract: Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous too

agentsarxiv-cs-ai
3 Jun 2026
Safety

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

DGX agent

arXiv:2602.08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving com

safetyarxiv-cs-ai
3 Jun 2026
Agents

Acting with AI: An Interaction-Based Framework for Agentic Tort Liability

DGX agent

arXiv:2606.00518v1 Announce Type: new Abstract: Agentic AI systems can plan over multiple steps, use tools, and execute tasks over time. When such systems cause harm, tort law struggles to allocate re

agentsarxiv-cs-ai
2 Jun 2026
Agents

Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows

DGX agent

arXiv:2602.14849v2 Announce Type: replace-cross Abstract: LLM agents execute multi-step workflows that mutate external state through tools. Common orchestrators treat tool return as the settlement tri

agentsarxiv-cs-ai
2 Jun 2026
Agents

Beyond One-shot: AI Agents for Learning in Field Experiments

DGX agent

arXiv:2606.02458v1 Announce Type: new Abstract: Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention d

agentsarxiv-cs-ai
2 Jun 2026
Agents

Can LLM Agents Sustain Long-Horizon Organizational Dynamics?

DGX agent

arXiv:2606.01199v1 Announce Type: new Abstract: Large language agents are increasingly used for social simulation, yet it remains unclear whether they can sustain coherent behavior in structured organ

agentsarxiv-cs-ai
2 Jun 2026
Agents

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

DGX agent

arXiv:2601.19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows tha

agentsarxiv-cs-ai
2 Jun 2026
Agents

'Do Not Mention This to the User': Detecting and Understanding Malicious Agent Skills

DGX agent

arXiv:2602.06547v3 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scrip

agentsarxiv-cs-ai
2 Jun 2026
Agents

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

DGX agent

arXiv:2606.01365v1 Announce Type: new Abstract: Tool-using multi-agent large language model (LLM) systems spend computation through model tokens, tool calls, retries, and code execution before produci

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

DGX agent

arXiv:2606.00660v1 Announce Type: new Abstract: Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a p

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools

DGX agent

arXiv:2606.02483v1 Announce Type: cross Abstract: Tool-augmented language agents speculatively issue likely future tool calls to hide latency, but those calls leak inferred user intent to external ser

agentsarxiv-cs-ai
2 Jun 2026
Safety

HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems

DGX agent

arXiv:2606.01779v1 Announce Type: new Abstract: LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent

safetyarxiv-cs-cl
2 Jun 2026
Agents

Iteris: Agentic Research Loops for Computational Mathematics

DGX agent

arXiv:2606.02484v1 Announce Type: new Abstract: Recent advances in large language models and agentic AI systems have enabled significant progress in mathematical discovery, from solving competition pr

agentsarxiv-cs-ai
2 Jun 2026
Safety

MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems

DGX agent

arXiv:2601.14230v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) are emerging as promising socio-collaborative companions for emotional and cognitive support. However, existing syst

safetyarxiv-cs-ai
2 Jun 2026
Agents

MemPro: Agentic Memory Systems as Evolvable Programs

DGX agent

arXiv:2606.00619v1 Announce Type: cross Abstract: Long-horizon autonomous agents require memory systems to retain historical information, track evolving states, and reuse relevant knowledge beyond fin

agentsarxiv-cs-ai
2 Jun 2026
Agents

Monitoring Agentic Systems Before They're Reliable

DGX agent

arXiv:2606.02494v1 Announce Type: cross Abstract: Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design

DGX agent

arXiv:2606.00555v1 Announce Type: new Abstract: Structure-based drug design increasingly employs LLM agents to iteratively refine ligands against a target pocket, yet a viable ligand must satisfy two

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Scaling Agentic Capabilities via Grounded Interaction Synthesis

DGX agent

arXiv:2606.02001v1 Announce Type: new Abstract: General agentic intelligence hinges on the ability to interact with diverse real-world tools to complete complex tasks, a capability fundamentally tied

agentsarxiv-cs-cl
2 Jun 2026
Agents

Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration

DGX agent

arXiv:2410.02511v2 Announce Type: replace Abstract: With expansive state-action spaces, efficient multi-agent exploration remains a longstanding challenge in reinforcement learning. Although pursuing

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

ForecastCompass: Guiding Agentic Forecasting with Adaptive Factor Memory

DGX agent

arXiv:2605.30858v1 Announce Type: new Abstract: Agentic forecasting is important for decision-making in dynamic environments, but it remains challenging because agents must reason from incomplete, tim

model-releasesarxiv-cs-lg
1 Jun 2026
Agents

Generalized Intention Modeling in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.31318v1 Announce Type: new Abstract: Modeling an opponent's intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learn

agentsarxiv-cs-lg
1 Jun 2026
Agents

HypoAgent: An Agentic Framework for Interactive Abductive Hypothesis Generation over Knowledge Graphs

DGX agent

arXiv:2605.31370v1 Announce Type: new Abstract: Abductive reasoning over knowledge graphs aims to generate logical hypotheses that explain observed entities or facts. Existing controllable hypothesis

agentsarxiv-cs-ai
1 Jun 2026
Model Releases

MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents

DGX agent

arXiv:2605.30727v1 Announce Type: new Abstract: Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external q

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Multi-Agent Teams Hold Experts Back

DGX agent

arXiv:2602.01011v4 Announce Type: replace-cross Abstract: Multi-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-spe

safetyarxiv-cs-ai
1 Jun 2026
← Previous
1…3940414243…233
Next →