AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
30 Jun 2026

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vecto…

Model ReleasesDGX agent

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vectors distilled from contrastive historical experience. Textual

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

AgentsDGX agent

arXiv:2606.29932v1 Announce Type: new Abstract: Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and

SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents

TutorialsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.28434v1 Announce Type: cross Abstract: Long-horizon software engineering agents often need to manage lengthy and noisy interaction histories under limited context budgets. Existing memory m

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing o…

Model ReleasesDGX agent

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing our journey and walking you through building your first agent

TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging

AgentsDGX agent

arXiv:2606.29763v1 Announce Type: cross Abstract: Topological data analysis (TDA), particularly persistent homology (PH), captures geometric structural properties in medical images (e.g., connected co

29 Jun 2026

Agent confidence on the technical frontier

AgentsDGX agent

Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mount

Agentic Episodic Control

AgentsDGX agent

arXiv:2506.01442v2 Announce Type: replace Abstract: Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attempt to alle

Agentic Hardware Design as Repository-Level Code Evolution

Model ReleasesDGX agent

arXiv:2606.28279v1 Announce Type: cross Abstract: We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled int

From Detection to Action: Using LLM Agents for Fault-Tolerant Control

Model ReleasesDGX agent

arXiv:2606.28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constr

Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs

SafetyDGX agent

arXiv:2601.21233v2 Announce Type: replace Abstract: Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-d

One day left to submit your projects for the Hermes Agent Accelerated Business Hackathon presented by @NVIDIAAI × @stripe × @NousResearch! S…

HardwareDGX agent

One day left to submit your projects for the Hermes Agent Accelerated Business Hackathon presented by @NVIDIAAI × @stripe × @NousResearch! Submissions close 11:59 PM PT tomorrow, June 30th. Last minut

Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems

SafetyDGX agent

arXiv:2505.23847v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint disas

Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents

Model ReleasesDGX agent

arXiv:2606.27472v1 Announce Type: cross Abstract: Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revise

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

Model ReleasesDGX agent

arXiv:2606.27669v1 Announce Type: new Abstract: Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval a

26 Jun 2026

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

Model ReleasesDGX agent

arXiv:2606.26216v1 Announce Type: cross Abstract: We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability det

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

AgentsDGX agent

arXiv:2606.21649v2 Announce Type: replace Abstract: Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Model ReleasesDGX agent

arXiv:2606.26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-doma

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

Model ReleasesDGX agent

arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tas

Prompt, Plan, Extract: Zero-Shot Agentic LLMs Workflows for Lung Pathology Extraction from Clinical Narratives

AgentsDGX agent

arXiv:2606.19852v2 Announce Type: replace Abstract: Information extraction from pathology reports is essential for cancer staging, tumor registry population. Yet key data remains embedded in narrative

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

AgentsDGX agent

arXiv:2606.26722v1 Announce Type: new Abstract: The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypot

25 Jun 2026

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

SafetyDGX agent

arXiv:2606.25421v1 Announce Type: new Abstract: Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. Howeve

GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2606.25073v1 Announce Type: new Abstract: In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch fo

MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models

AgentsDGX agent

arXiv:2606.24815v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging robotic control systems that integrate perception, language understanding, and action generation in a

Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

AgentsDGX agent

arXiv:2606.25361v1 Announce Type: new Abstract: Prior research on memory mechanism in RAG-based conversational system has emphasized how memory is stored and retrieved. However, far less is known abou

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse

SafetyDGX agent

arXiv:2606.25389v1 Announce Type: new Abstract: Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings

Patronus AI, which builds simulated digital environments for evaluating AI agents, raised a 50M Series B led by Greenfield, bringing its total funding to 70M (Marina Temkin/TechCrunch)

IndustryDGX agent

Marina Temkin / TechCrunch: Patronus AI, which builds simulated digital environments for evaluating AI agents, raised a 50M Series B led by Greenfield, bringing its total funding to 70M — AI agents ar

24 Jun 2026

Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems

SafetyDGX agent

arXiv:2606.24416v1 Announce Type: new Abstract: Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives a

AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

Model ReleasesDGX agent

arXiv:2606.24526v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grou

ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.24601v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives.

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

Model ReleasesDGX agent

arXiv:2606.17698v2 Announce Type: replace Abstract: As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the que

Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

SafetyDGX agent

arXiv:2606.24010v1 Announce Type: new Abstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approach

Varying Bundle Size Reactive Multi-Task Assignment using Selective Cost Estimation for Multi-Agent Systems

AgentsDGX agent

arXiv:2606.24462v1 Announce Type: new Abstract: This paper presents a scalable framework for multi-robot task allocation in complex environments where estimating task execution costs is computationall

23 Jun 2026

Hypothesis-Disciplined Multi-Agent Automated Formalization of Asymptotic Statistical Theory

AgentsDGX agent

arXiv:2606.20642v1 Announce Type: cross Abstract: Asymptotic statistical theory is a challenging domain for AI-assisted formalization: its central results mix convergence statements, asymptotic expans

Log Analytics is now Observability Analytics: Query logs and traces with SQL

AgentsDGX agent

To effectively operate and troubleshoot applications, developers and site reliability engineers (SREs) need to understand the full context of their system's behavior, typically as part of their loggin

RAPID: A Reproducible Multi-Agent Pipeline for Interpretable Disaster Damage Assessment from Satellite and Street-View Imagery

AgentsDGX agent

arXiv:2606.21819v1 Announce Type: new Abstract: Due to the increasing frequency and intensity of extreme climate events, there is a clear demand for intelligent, scalable, and autonomous approaches to

SPARC: A Multi-Agent System for Electrical Circuit Question Answering

AgentsDGX agent

arXiv:2606.20643v1 Announce Type: cross Abstract: Electrical circuit diagram QA tasks require complex mathematical reasoning, which remains challenging for multimodal LLMs. We present SPARC, a multi-a

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams

Model ReleasesDGX agent

arXiv:2606.20629v1 Announce Type: cross Abstract: LLM agents are increasingly deployed as multi-role teams, where tasks are divided across specialized roles such as planner, executor, and verifier. In

WebCryptoAgent: Agentic Crypto Trading with Web Informatics

AgentsDGX agent

arXiv:2601.04687v2 Announce Type: replace Abstract: Cryptocurrency trading increasingly depends on timely integration of heterogeneous web information and market microstructure signals to support shor

22 Jun 2026

Ai2 just released TMax 27B on Hugging Face A 27B terminal agent that hits 42.7% on Terminal Bench 2.0, rivaling models 40× its size.

Model ReleasesDGX agent

AI2 released TMax 27B, a 27 billion parameter terminal agent model available on Hugging Face that achieves 42.7% performance on Terminal Bench 2.0, matching the capabilities of much larger models desp

Join us in our Discord this time tomorrow for Office Hours! Our topic will be the @NVIDIAAI × @stripe × @NousResearch Hermes Agent Accelerat…

HardwareDGX agent

Join us in our Discord this time tomorrow for Office Hours! Our topic will be the @NVIDIAAI × @stripe × @NousResearch Hermes Agent Accelerated Business Hackathon which concludes at the end of the mont

21 Jun 2026

Temporary Cloudflare Accounts for AI agents

Model ReleasesDGX agent

Temporary Cloudflare Accounts for AI agents The announcement says this is 'for AI agents' but (as is pretty common these days) the AI hook isn't really necessary, this is an interesting feature for ev

11 Jun 2026

A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents

Model ReleasesDGX agent

arXiv:2606.12320v1 Announce Type: new Abstract: Enterprise security was built to govern data boundaries: the protected surface was data at rest and in transit, and the controls -- access control, data

Agent Skill Evaluation and Evolution: Frameworks and Benchmarks

Model ReleasesDGX agent

arXiv:2606.11435v1 Announce Type: new Abstract: The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evalua

Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment

Local AiDGX agent

arXiv:2606.11672v1 Announce Type: cross Abstract: This paper explores the value of agentic AI tools for cybersecurity purposes. We evaluate the efficacy of a general-purpose GenAI Large Language Model

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

Model ReleasesDGX agent

arXiv:2606.12344v1 Announce Type: cross Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-ben

Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies

SafetyDGX agent

arXiv:2602.18291v2 Announce Type: replace Abstract: Online Multi-Agent Reinforcement Learning (MARL) is a prominent framework for efficient agent coordination. Crucially, enhancing policy expressivene

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems

AgentsDGX agent

arXiv:2606.12236v1 Announce Type: cross Abstract: Many autonomous driving systems are increasingly incorporating foundation models to improve generalization and handle long-tail scenarios. However, th

Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA

AgentsDGX agent

arXiv:2603.08501v3 Announce Type: replace Abstract: Large language models (LLMs) can answer religious knowledge queries fluently, yet they often hallucinate and misattribute sources, which is especial

Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark

Model ReleasesDGX agent

arXiv:2602.19502v2 Announce Type: replace Abstract: Agentic AI systems are increasingly capable of autonomous data science workflows, yet clinical prediction tasks demand domain expertise that purely

INFRAMIND: Infrastructure-Aware Multi-Agent Orchestration

HardwareDGX agent

arXiv:2606.11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and mo

MARIC: Multi-Agent Reasoning for Image Classification

Model ReleasesDGX agent

arXiv:2509.14860v2 Announce Type: replace-cross Abstract: Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fi

MedCTA: A Benchmark for Clinical Tool Agents

Model ReleasesDGX agent

arXiv:2606.11702v1 Announce Type: cross Abstract: To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acqui

PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents

Local AiDGX agent

arXiv:2606.12329v1 Announce Type: new Abstract: AI coding assistants now support a growing share of software work, from quick scripts to production applications. Yet these agents remain largely statel

Sustainability assessment using multimodal AI agents

AgentsDGX agent

arXiv:2507.17012v2 Announce Type: replace Abstract: Reducing the rapidly growing environmental impact of the computing industry requires assessing the emissions of electronics at scale. However, a tra

10 Jun 2026

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Model ReleasesDGX agent

arXiv:2606.11150v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experime

Assessing Automated Prompt Injection Attacks in Agentic Environments

Model ReleasesDGX agent

arXiv:2606.10525v1 Announce Type: cross Abstract: Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effec

au-Rec: A Verifiable Benchmark for Agentic Recommender Systems

Model ReleasesDGX agent

arXiv:2606.10156v1 Announce Type: cross Abstract: As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benc

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

AgentsDGX agent

arXiv:2510.04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-tho

Deployment-Time Memorization in Foundation-Model Agents

Model ReleasesDGX agent

arXiv:2606.10062v1 Announce Type: new Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time fun

everything your agents do, in a bucket you own. private by default. session traces, task claims, artifacts — claude code, hermes, whatever y…

Model ReleasesDGX agent

everything your agents do, in a bucket you own. private by default. session traces, task claims, artifacts — claude code, hermes, whatever you run. and make your agents friends via messaging https://h

← Previous
1…9394959697…300
Next →