AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Agents

Prompt-to-Paper: Agentic AI System for Bioinformatics

DGX agent

arXiv:2607.05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical defi

agentsarxiv-cs-ai
8 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

DGX agent

arXiv:2607.06411v1 Announce Type: cross Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

DGX agent

arXiv:2606.13757v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is mer

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

DGX agent

arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their nav

model-releasesarxiv-cs-cv
8 Jul 2026
Agents

Agentic Very Long Video Understanding

DGX agent

arXiv:2601.18157v3 Announce Type: replace Abstract: The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual underst

agentsarxiv-cs-cv
7 Jul 2026
Agents

AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression

DGX agent

arXiv:2512.13956v4 Announce Type: replace-cross Abstract: Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and met

agentsarxiv-cs-ai
7 Jul 2026
Model Releases

AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation

DGX agent

arXiv:2607.02520v1 Announce Type: cross Abstract: Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether gener

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

EEG-SpikeAgent: Agentic Closed-Loop Program Synthesis for Automated EEG Spike Detection

DGX agent

arXiv:2607.04558v1 Announce Type: cross Abstract: Automated detection of interictal epileptiform discharges in scalp electroencephalography (EEG) is clinically important, but recent high-performing de

agentsarxiv-cs-ai
7 Jul 2026
Agents

Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning

DGX agent

arXiv:2607.02588v1 Announce Type: cross Abstract: Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally

agentsarxiv-cs-ai
7 Jul 2026
Model Releases

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

DGX agent

arXiv:2606.20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents

DGX agent

arXiv:2607.04089v1 Announce Type: new Abstract: Lifelong agents need more than larger context windows and better retrieval. They need memories that can persist, evolve, and be corrected without forcin

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe

DGX agent

arXiv:2607.03451v1 Announce Type: cross Abstract: While skill optimization for autonomous agents has gained traction, existing methods rely on complex pipelines. This leaves a fundamental question una

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

CausalSteward: An Agentic Divide-Conquer-Combine Copilot for Causal Discovery

DGX agent

arXiv:2607.01936v1 Announce Type: cross Abstract: Learning causal models from high-dimensional data is a significant challenge, particularly in real-world settings where violations of core assumptions

agentsarxiv-cs-ai
3 Jul 2026
Model Releases

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

DGX agent

arXiv:2606.14249v2 Announce Type: replace Abstract: AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model obs

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

DGX agent

arXiv:2604.04532v2 Announce Type: replace-cross Abstract: Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

DGX agent

arXiv:2607.01084v1 Announce Type: new Abstract: While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynami

model-releasesarxiv-cs-ai
2 Jul 2026
Agents

Coachable agents for interactive gameplay

DGX agent

arXiv:2607.00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

DGX agent

arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -

model-releasesarxiv-cs-cl
2 Jul 2026
Agents

Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization

DGX agent

arXiv:2603.26270v2 Announce Type: replace-cross Abstract: Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

DGX agent

arXiv:2607.01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. How

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

DGX agent

arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestr

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents

DGX agent

arXiv:2606.30970v1 Announce Type: new Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

DGX agent

arXiv:2606.31200v1 Announce Type: new Abstract: Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based me

model-releasesarxiv-cs-ai
1 Jul 2026
Agents

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

DGX agent

arXiv:2606.31134v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evad

agentsarxiv-cs-ai
1 Jul 2026
Model Releases

Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

DGX agent

arXiv:2606.31371v1 Announce Type: cross Abstract: When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

Verification-Gated Agentic Mission-State Governance for Intelligent Industrial Multi-Robot Systems

DGX agent

arXiv:2606.31339v1 Announce Type: new Abstract: Agentic artificial intelligence is increasingly used to decompose industrial tasks, propose robot actions, and adapt execution plans in dynamic cyber-ph

safetyarxiv-cs-ro
1 Jul 2026
Model Releases

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

DGX agent

arXiv:2606.29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observa

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation

DGX agent

arXiv:2606.28896v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) data augmentation is important for improving the generalization of data-driven SAR interpretation models, yet practical

agentsarxiv-cs-ai
30 Jun 2026
Safety

Agent Safety Is Action Alignment

DGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

safetyarxiv-cs-ai
30 Jun 2026
Agents

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

DGX agent

arXiv:2606.17555v2 Announce Type: replace-cross Abstract: Banks face two threat families with fundamentally different detection requirements: signature-based fraud (card-not-present attacks, account t

agentsarxiv-cs-ai
30 Jun 2026
Safety

Analytic Concept-Centric Memory for Agentic Embodied Manipulation

DGX agent

arXiv:2606.29774v1 Announce Type: new Abstract: Long-horizon embodied manipulation requires agents to remember persistent objects, track changing scene states, and reuse prior interaction knowledge. H

safetyarxiv-cs-ro
30 Jun 2026
Agents

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

DGX agent

arXiv:2606.29746v1 Announce Type: new Abstract: Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains

agentsarxiv-cs-ai
30 Jun 2026
Safety

Entity Binding Failures in Tool-Augmented Agents

DGX agent

arXiv:2606.30531v1 Announce Type: new Abstract: Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requeste

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard

DGX agent

arXiv:2606.30005v1 Announce Type: new Abstract: Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management age

model-releasesarxiv-cs-cl
30 Jun 2026
Agents

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

DGX agent

arXiv:2606.29932v1 Announce Type: new Abstract: Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and

agentsarxiv-cs-ai
30 Jun 2026
Tutorials

SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents

DGX agent

arXiv:2606.28434v1 Announce Type: cross Abstract: Long-horizon software engineering agents often need to manage lengthy and noisy interaction histories under limited context budgets. Existing memory m

tutorialsarxiv-cs-ai
30 Jun 2026
Agents

TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging

DGX agent

arXiv:2606.29763v1 Announce Type: cross Abstract: Topological data analysis (TDA), particularly persistent homology (PH), captures geometric structural properties in medical images (e.g., connected co

agentsarxiv-cs-ai
30 Jun 2026
Agents

Agentic Episodic Control

DGX agent

arXiv:2506.01442v2 Announce Type: replace Abstract: Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attempt to alle

agentsarxiv-cs-ai
29 Jun 2026
Model Releases

Agentic Hardware Design as Repository-Level Code Evolution

DGX agent

arXiv:2606.28279v1 Announce Type: cross Abstract: We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled int

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

From Detection to Action: Using LLM Agents for Fault-Tolerant Control

DGX agent

arXiv:2606.28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constr

model-releasesarxiv-cs-lg
29 Jun 2026
Safety

Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs

DGX agent

arXiv:2601.21233v2 Announce Type: replace Abstract: Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-d

safetyarxiv-cs-ai
29 Jun 2026
Safety

Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems

DGX agent

arXiv:2505.23847v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint disas

safetyarxiv-cs-ai
29 Jun 2026
Model Releases

Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents

DGX agent

arXiv:2606.27472v1 Announce Type: cross Abstract: Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revise

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

DGX agent

arXiv:2606.27669v1 Announce Type: new Abstract: Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval a

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

DGX agent

arXiv:2606.26216v1 Announce Type: cross Abstract: We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability det

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

DGX agent

arXiv:2606.21649v2 Announce Type: replace Abstract: Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This

agentsarxiv-cs-cl
26 Jun 2026
Model Releases

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

DGX agent

arXiv:2606.26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-doma

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

DGX agent

arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tas

model-releasesarxiv-cs-ai
26 Jun 2026
← Previous
1…6667686970…236
Next →