AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
7 May 2026

Anthropic is letting Claude agents ‘dream’ so they don’t sleep on the job

Model ReleasesDGX agent

Anthropic PBC said today it’s giving its AI agents the ability to “dream” and remember past interactions and work they’ve performed so they can identify recurring mistakes and improve over time. In an

Closed-Loop Vision-Language Planning for Multi-Agent Coordination

Model ReleasesDGX agent

arXiv:2502.10148v3 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) struggles with sample efficiency, interpretability, and generalization. While Large Language M

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

GeoDecider: A Coarse-to-Fine Agentic Workflow for Explainable Lithology Classification

AgentsDGX agent

arXiv:2605.03383v1 Announce Type: new Abstract: Lithology classification aims to infer subsurface rock types from well-logging signals, supporting downstream applications like reservoir characterizati

Multi-Agent Strategic Games with LLMs

AgentsDGX agent

arXiv:2605.03604v1 Announce Type: cross Abstract: This paper asks whether large language models (LLMs) can be used to study the strategic foundations of conflict and cooperation. I introduce LLMs as e

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

Model ReleasesDGX agent

arXiv:2605.04019v1 Announce Type: new Abstract: AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a

Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense

Model ReleasesDGX agent

arXiv:2605.03034v1 Announce Type: new Abstract: Agentic systems involved in high-stake decision-making under adversarial pressure need formal guarantees not offered by existing approaches. Motivated b

6 May 2026

AI Agents for Sustainable SMEs: A Green ESG Assessment Framework

AgentsDGX agent

arXiv:2605.00841v1 Announce Type: new Abstract: This study presents a novel, AI-driven framework for assessing Environmental, Social, and Governance (ESG) performance in European small and medium-size

Architectural Obsolescence of Unhardened Agentic-AI Runtimes

SafetyDGX agent

arXiv:2605.01740v1 Announce Type: cross Abstract: An agentic-AI runtime issues tool calls, sends messages, and actuates devices on behalf of an LLM. Catching the four ways an action can diverge from i

GDPR Auto-Formalization with AI Agents and Human Verification

AgentsDGX agent

arXiv:2604.14607v2 Announce Type: replace Abstract: We study the overall process of automatic formalization of GDPR provisions using large language models, within a human-in-the-loop verification fram

LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

Model ReleasesDGX agent

arXiv:2605.01394v1 Announce Type: cross Abstract: Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Alth

MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration

SafetyDGX agent

arXiv:2605.01507v1 Announce Type: new Abstract: Prior studies report that partial driving automation can increase the cognitive demands on human drivers. This effect largely arises from human drivers'

Practical Limits of Autonomous Test Repair: A Multi-Agent Case Study with LLM-Driven Discovery and Self-Correction

AgentsDGX agent

arXiv:2605.01471v1 Announce Type: cross Abstract: Maintaining reliable UI test suites in large-scale enterprise applications is a persistent and costly challenge. We present an industrial case study o

The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate

Model ReleasesDGX agent

arXiv:2605.00914v1 Announce Type: cross Abstract: Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

Model ReleasesDGX agent

arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte

When Alignment Isn't Enough: Response-Path Attacks on LLM Agents

Model ReleasesDGX agent

arXiv:2605.02187v1 Announce Type: cross Abstract: Bring-Your-Own-Key (BYOK) agent architectures let users route LLM traffic through third-party relays, creating a critical integrity gap: a malicious r

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.02463v1 Announce Type: cross Abstract: Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, t

5 May 2026

Autonomous QA Agent: A Retrieval-Augmented Framework for Reliable Selenium Script Generation

AgentsDGX agent

arXiv:2601.06034v1 Announce Type: cross Abstract: Software testing is critical in the software development lifecycle, yet translating requirements into executable test scripts remains manual and error

CP-SynC: Multi-Agent Zero-Shot Constraint Modeling in MiniZinc with Synthesized Checkers

Model ReleasesDGX agent

arXiv:2605.01675v1 Announce Type: cross Abstract: Constraint Programming (CP) is a powerful paradigm for solving combinatorial problems, yet translating natural language problem descriptions into exec

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture

Model ReleasesDGX agent

arXiv:2605.01567v1 Announce Type: cross Abstract: Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineeri

GA-VisAgent: A Multi-Agent application for code generation and visualization in interactive learning

AgentsDGX agent

arXiv:2605.01299v1 Announce Type: new Abstract: Geometric Algebra (GA) presents challenges to learners due to its highly abstract mathematical structure and complex operational rules, as translating a

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

Model ReleasesDGX agent

arXiv:2505.10887v3 Announce Type: replace Abstract: This paper introduces extsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, i

LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery

AgentsDGX agent

arXiv:2508.12232v3 Announce Type: replace-cross Abstract: Issue-to-commit link recovery in software repositories is fundamental to software traceability and project management, yet it remains a challe

MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding

Model ReleasesDGX agent

arXiv:2510.08804v3 Announce Type: replace Abstract: We present MOSAIC, a multi-agent Large Language Model (LLM) framework for solving challenging scientific coding tasks. Unlike general-purpose coding

Planner Matters! An Efficient and Unbalanced Multi-agent Collaboration Framework for Long-horizon Planning

Model ReleasesDGX agent

arXiv:2605.02168v1 Announce Type: cross Abstract: Language model (LM)-based agents have demonstrated promising capabilities in automating complex tasks from natural language instructions, yet they con

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

SafetyDGX agent

arXiv:2605.01416v1 Announce Type: cross Abstract: The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy

4 May 2026

Pinecone Nexus: The Knowledge Engine for Agents

ToolsDGX agent

Pinecone Nexus is a knowledge infrastructure platform designed to enable AI agents to access, retrieve, and reason over enterprise data at scale. The system addresses the challenge of integrating larg

2 May 2026

[AINews] AI Engineer World's Fair — Autoresearch, Memory, World Models, Tokenmaxxing, Agentic Commerce, and Vertical AI Call for Speakers

AgentsDGX agent

The AI Engineer World's Fair is a conference event featuring talks on emerging AI research and engineering topics including autoresearch systems, advanced memory architectures, world models, tokenizat

1 May 2026

AgentEconomist: An End-to-end Agentic System Translating Economic Intuitions into Executable Computational Experiments

AgentsDGX agent

arXiv:2604.27725v1 Announce Type: cross Abstract: A long-standing challenge in economics lies not in the lack of intuition, but in the difficulty of translating intuitive insights into verifiable rese

Context as Prior: Bayesian-Inspired Intent Inference for Non-Speaking Agents with a Household Cat Testbed

ApplicationsDGX agent

arXiv:2604.27445v1 Announce Type: new Abstract: Many agents in real-world environments cannot reliably communicate their goals through language, including household pets, pre-verbal infants, and other

DeepTutor: Towards Agentic Personalized Tutoring

Model ReleasesDGX agent

arXiv:2604.26962v1 Announce Type: cross Abstract: Education represents one of the most promising real-world applications for Large Language Models (LLMs). However, conventional tutoring systems rely o

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

SafetyDGX agent

arXiv:2604.27045v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memor

End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians

AgentsDGX agent

arXiv:2604.27309v1 Announce Type: new Abstract: Clinical AI systems require not just point-in-time evaluation but continuous governance: the ongoing practice of monitoring, evaluating, iterating, and

Heterogeneous Scientific Foundation Model Collaboration

AgentsDGX agent

arXiv:2604.27351v1 Announce Type: new Abstract: Agentic large language model systems have demonstrated strong capabilities. However, their reliance on language as the universal interface fundamentally

Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents

SafetyDGX agent

arXiv:2604.27283v1 Announce Type: cross Abstract: Large language model (LLM)-based coding agents increasingly rely on external memory to reuse prior debugging experience, repair traces, and repository

MCP vs. CLI Skills for agents: what our eval found (and which you should use)

AgentsDGX agent

Twitter said pick a side. The eval said the question was wrong. Six months ago, MCP (model context protocol) was the hot new thing: tool usage with a built-in discovery... The post MCP vs. CLI Skills

Optimal Stop-Loss and Take-Profit Parameterization for Autonomous Trading Agent Swarm

AgentsDGX agent

arXiv:2604.27150v1 Announce Type: new Abstract: Autonomous crypto trading systems often spend most of their design effort on finding entries, while exits are left to fixed rules that are rarely tested

Progressive Multi-Agent Reasoning for Biological Perturbation Prediction

Model ReleasesDGX agent

arXiv:2602.07408v2 Announce Type: replace Abstract: Predicting gene regulation responses to biological perturbations requires reasoning about underlying biological causalities. While large language mo

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

Model ReleasesDGX agent

arXiv:2604.28093v1 Announce Type: new Abstract: Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the

30 Apr 2026

Agentic opportunity: OpenAI and Stripe build for rising tide of new entrepreneurial firms

AgentsDGX agent

In corporate parlance, any metric on a chart that shows rapid, significant growth over a short period of time is known as a “hockey stick.” On Wednesday, Stripe Inc. Chief Executive Patrick Collison s

Creating highly efficient agents: 450M tool-calling tokens distilled for post-training from top open-source models

Model ReleasesDGX agent

Harnesses If you've used Claude Code or Codex, you've used a harness. A harness is the infrastructure layer that wraps an AI coding agent and decides how it operates, what it can touch, and how you me

Evaluating Strategic Reasoning in Forecasting Agents

SafetyDGX agent

arXiv:2604.26106v1 Announce Type: new Abstract: Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to t

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents

Model ReleasesDGX agent

arXiv:2511.02399v2 Announce Type: replace-cross Abstract: Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requireme

Lightweight Quantum Agent for Edge Systems: Joint PQC and NOMA Resource Allocation

AgentsDGX agent

arXiv:2604.25980v1 Announce Type: cross Abstract: In the context of quantum secure scenarios, existing research on mobile edge devices and intelligent computing and edge (ICE) systems based on the Non

Provable Coordination for LLM Agents via Message Sequence Charts

Local AiDGX agent

arXiv:2604.17612v2 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are difficult to reason about. Coordination errors such as deadlocks or type-mismatc

StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

AgentsDGX agent

arXiv:2508.01875v4 Announce Type: replace Abstract: Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offl

29 Apr 2026

Auvik launches Aurora AI agents to speed ticket resolution and prevent outages

Model ReleasesDGX agent

Information technology management software provider Auvik Networks Inc. today announced the launch of Auvik Aurora: artificial intelligence-powered IT agents that are designed to help IT professionals

Aviatrix launches AI agent containment platform for cloud workloads

Model ReleasesDGX agent

Aviatrix Inc. today announced the launch of a new platform designed to contain artificial intelligence agents and enforce security controls and communications across AI workloads without changing AI a

Databricks and Stripe Projects: Infrastructure Built for Agents

ApplicationsDGX agent

Databricks and Stripe have collaborated on infrastructure projects designed to support AI agents, likely focusing on data processing and payment integration capabilities. The initiative demonstrates h

Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver

Model ReleasesDGX agent

arXiv:2604.25067v1 Announce Type: cross Abstract: Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks mea

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

Model ReleasesDGX agent

arXiv:2604.24929v1 Announce Type: new Abstract: Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editin

28 Apr 2026

Evaluating the Search Agent in a Parallel World

Model ReleasesDGX agent

arXiv:2603.04751v2 Announce Type: replace Abstract: Integrating web search tools has significantly extended the capability of LLMs to address open-world, real-time, and long-tail problems. However, ev

GradMAP: Gradient-Based Multi-Agent Proximal Learning for Grid-Edge Flexibility

Model ReleasesDGX agent

arXiv:2604.24549v1 Announce Type: cross Abstract: Coordinating large populations of grid-edge devices requires learning methods that remain fully decentralised in deployment while still respecting thr

InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning

Model ReleasesDGX agent

arXiv:2508.19679v2 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have enabled mobile agents to perceive and interact with real-world mobile environments based on hu

Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training

Model ReleasesDGX agent

arXiv:2506.20332v4 Announce Type: replace Abstract: Vision-language model-based mobile agents have gained the ability to understand complex instructions and mobile screenshots, benefiting from reinfor

Scalable Agentic Reasoning for Designing Biologics Targeting Intrinsically Disordered Proteins

Model ReleasesDGX agent

arXiv:2512.15930v2 Announce Type: replace-cross Abstract: Intrinsically disordered proteins (IDPs) represent crucial therapeutic targets due to their significant role in disease -- approximately 80% o

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

Model ReleasesDGX agent

arXiv:2509.13021v2 Announce Type: replace-cross Abstract: This work introduces xOffense, an AI-driven, multi-agent penetration testing framework that shifts the process from labor-intensive, expert-dr

27 Apr 2026

A Probabilistic Framework for Hierarchical Goal Recognition

AgentsDGX agent

arXiv:2604.22256v1 Announce Type: cross Abstract: Goal recognition aims to infer an agent's goal from observations of its behaviour. In realistic settings, recognition can benefit from exploiting hier

Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning

AgentsDGX agent

arXiv:2604.22191v1 Announce Type: cross Abstract: In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a

Happy to announce that Hermes Agent's repo just surpassed Anthropic's Claude Code repo

Model ReleasesDGX agent

Nous Research announced that their Hermes Agent repository has surpassed Anthropic's Claude Code repository in popularity metrics, likely referring to GitHub stars or similar engagement measures. This

← Previous
1…109110111112113…300
Next →