AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Agents

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

DGX agent

arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal ev

agentsarxiv-cs-cv
31 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

DGX agent

arXiv:2607.24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observa

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

DGX agent

arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context,

model-releasesarxiv-cs-ai
29 Jul 2026
Agents

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

DGX agent

arXiv:2607.25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. H

agentsarxiv-cs-ai
29 Jul 2026
Safety

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

DGX agent

arXiv:2607.25297v1 Announce Type: cross Abstract: The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactio

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

DGX agent

arXiv:2607.25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Pri

model-releasesarxiv-cs-ai
29 Jul 2026
Agents

Towards an Agent Operating System - Lessons from Classical and Cloud OS

DGX agent

arXiv:2607.25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, f

agentsarxiv-cs-ai
29 Jul 2026
Safety

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

DGX agent

arXiv:2607.22724v1 Announce Type: cross Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing traject

safetyarxiv-cs-ai
28 Jul 2026
Agents

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

DGX agent

arXiv:2607.22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve dir

agentsarxiv-cs-ai
28 Jul 2026
Model Releases

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

DGX agent

arXiv:2607.22732v1 Announce Type: new Abstract: LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

DGX agent

arXiv:2607.22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. H

safetyarxiv-cs-ai
28 Jul 2026
Agents

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning

DGX agent

arXiv:2509.06278v4 Announce Type: replace Abstract: Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent large lang

agentsarxiv-cs-ai
28 Jul 2026
Agents

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

DGX agent

arXiv:2607.23734v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relayin

agentsarxiv-cs-lg
28 Jul 2026
Agents

Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach

DGX agent

arXiv:2601.17303v2 Announce Type: replace Abstract: As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce l

agentsarxiv-cs-lg
27 Jul 2026
Model Releases

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

DGX agent

arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contaminati

model-releasesarxiv-cs-cl
27 Jul 2026
Safety

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

DGX agent

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical align

safetyarxiv-cs-cl
27 Jul 2026
Model Releases

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

DGX agent

arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose m

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

DGX agent

arXiv:2607.20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score th

model-releasesarxiv-cs-ai
24 Jul 2026
Local Ai

ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models

DGX agent

arXiv:2607.20499v1 Announce Type: new Abstract: Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability. We presen

local-aiarxiv-cs-ai
24 Jul 2026
Agents

HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices

DGX agent

arXiv:2607.21019v1 Announce Type: new Abstract: Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisat

agentsarxiv-cs-ai
24 Jul 2026
Model Releases

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

DGX agent

arXiv:2607.11149v3 Announce Type: replace Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including lo

model-releasesarxiv-cs-ai
24 Jul 2026
Local Ai

Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

DGX agent

arXiv:2607.21495v1 Announce Type: new Abstract: AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments.

local-aiarxiv-cs-ai
24 Jul 2026
Agents

Agent-Centric Animal Pose Forecasting

DGX agent

arXiv:2607.19548v1 Announce Type: new Abstract: Understanding animal behavior at an algorithmic level -- what animals attend to, how they form internal models and plans, and how this maps to action --

agentsarxiv-cs-lg
23 Jul 2026
Model Releases

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

DGX agent

arXiv:2607.15434v3 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome

model-releasesarxiv-cs-ai
23 Jul 2026
Agents

NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem

DGX agent

arXiv:2607.19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology. We

agentsarxiv-cs-lg
23 Jul 2026
Agents

Personalized Recommendation Tool Learning via Autonomous Language Agents

DGX agent

arXiv:2607.19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive wo

agentsarxiv-cs-ai
23 Jul 2026
Model Releases

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

DGX agent

arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the fin

model-releasesarxiv-cs-ai
16 Jul 2026
Agents

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

DGX agent

arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior

agentsarxiv-cs-ai
16 Jul 2026
Model Releases

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

DGX agent

arXiv:2607.12894v1 Announce Type: new Abstract: Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, a

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

DGX agent

arXiv:2607.10079v2 Announce Type: replace Abstract: Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming

DGX agent

arXiv:2607.11913v1 Announce Type: cross Abstract: Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-li

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

DGX agent

arXiv:2607.08652v1 Announce Type: new Abstract: Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This pa

model-releasesarxiv-cs-ai
10 Jul 2026
Agents

GitLake: Git-for-data for the agentic lakehouse

DGX agent

arXiv:2607.08319v1 Announce Type: cross Abstract: We present GitLake, a Git-for-data design for an agent-first lakehouse. The system lifts single-table Iceberg snapshots into lakehouse-wide commits, b

agentsarxiv-cs-ai
10 Jul 2026
Safety

Agentic Data Environments

DGX agent

arXiv:2607.07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. Th

safetyarxiv-cs-ai
9 Jul 2026
Model Releases

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

DGX agent

arXiv:2607.07695v1 Announce Type: new Abstract: We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task

model-releasesarxiv-cs-ai
9 Jul 2026
Agents

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

DGX agent

arXiv:2607.07676v1 Announce Type: new Abstract: Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs n

agentsarxiv-cs-ai
9 Jul 2026
Safety

A toy framework for single and multi-agent human-AI curiosity ecosystems

DGX agent

arXiv:2607.06214v1 Announce Type: new Abstract: This paper offers a toy framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when, and why

safetyarxiv-cs-ai
8 Jul 2026
Agents

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

DGX agent

arXiv:2607.05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce A

agentsarxiv-cs-ai
8 Jul 2026
Agents

Delay-Aware Active Triangulation with Uncertainty-Driven Multi-Agent Reinforcement Learning for Counter-UAS

DGX agent

arXiv:2607.05957v1 Announce Type: new Abstract: Multi-agent active visual triangulation enables precise 3D localization of aerial targets by coordinating mobile observers with controllable cameras. Ho

agentsarxiv-cs-ro
8 Jul 2026
Model Releases

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

DGX agent

arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external env

model-releasesarxiv-cs-ai
8 Jul 2026
Agents

Agentic AI-RAN: Enabling Intent-Driven, Explainable and Self-Evolving Open RAN Intelligence

DGX agent

arXiv:2602.24115v2 Announce Type: replace Abstract: Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder t

agentsarxiv-cs-lg
7 Jul 2026
Agents

An Exploration of Agentic Information Fusion for Test Maintenance Prediction

DGX agent

arXiv:2607.04786v1 Announce Type: cross Abstract: Test maintenance is a critical, yet costly, activity - particularly as codebases rapidly evolve. To assist, we present MAST, a multi-agent framework t

agentsarxiv-cs-ai
7 Jul 2026
Model Releases

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games

DGX agent

arXiv:2607.04293v1 Announce Type: cross Abstract: Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally reli

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality

DGX agent

arXiv:2607.03691v1 Announce Type: cross Abstract: Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middlewa

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents

DGX agent

arXiv:2607.04718v1 Announce Type: new Abstract: Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This work

agentsarxiv-cs-ai
7 Jul 2026
Agents

kAgent: An execution-guided crash resolution agent for the Linux kernel

DGX agent

arXiv:2504.20412v3 Announce Type: replace-cross Abstract: Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive. However

agentsarxiv-cs-ai
7 Jul 2026
Safety

OpenTinker: Separating Concerns in Agentic Reinforcement Learning

DGX agent

arXiv:2601.07376v2 Announce Type: replace Abstract: We introduce extsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over shared exe

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration

DGX agent

arXiv:2607.03730v1 Announce Type: new Abstract: Conversational agents are increasingly embedded in human collaborative work, yet they remain fundamentally passive and reactive: they respond to explici

model-releasesarxiv-cs-cl
7 Jul 2026
← Previous
1…4546474849…233
Next →