AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
22 May 2026

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

Model ReleasesDGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2507.03674v3 Announce Type: replace Abstract: Extracting structured information from scientific literature is critical for accelerating discovery, yet Large Language Models (LLMs) often struggle

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

SafetyDGX agent

arXiv:2605.05765v2 Announce Type: replace Abstract: Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive intera

21 May 2026

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

Model ReleasesDGX agent

arXiv:2605.20530v1 Announce Type: cross Abstract: Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but the benchmarks used to evalu

Cool release from @p0. I think this use-case (agents paying content creators for access via @mpp) will be very big. Micropayment walls haven…

IndustryDGX agent

Cool release from @p0. I think this use-case (agents paying content creators for access via @mpp) will be very big. Micropayment walls haven't worked (as Clay Shirky anticipated many years ago) becaus

Creé mi propio agente de IA local centrado en Ollama y la ejecución de tareas prácticas.

Local AiDGX agent

This post describes a user's experience building a local AI agent using Ollama, an open-source tool for running large language models locally. The article likely covers the practical implementation st

Evolutionary Generation of Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2602.06511v3 Announce Type: replace Abstract: Large language model (LLM)-based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but design

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents

SafetyDGX agent

arXiv:2605.20246v1 Announce Type: new Abstract: Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple

Strong demand for agentic AI offerings helps Workday beat expectations and boost its stock

AgentsDGX agent

Shares of the financial services and human resources software company Workday Inc. rose more than 10% in late trading after it reported first-quarter earnings and revenue that came in higher than expe

Training Language Agents to Learn from Experience

Model ReleasesDGX agent

arXiv:2605.20477v1 Announce Type: cross Abstract: Language agents can adapt from experience in interactive environments, but current reflection-based methods can only self-correct within a single task

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

Model ReleasesDGX agent

arXiv:2605.20306v1 Announce Type: new Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous

Workday brings AI agents to IT service management and travel

ApplicationsDGX agent

Enterprise management software company Workday Inc. is expanding beyond core human resources and finance with new artificial intelligence agents for information technology service management and trave

Zscaler acquires Symmetry Systems to extend its AI agent security capabilities

AgentsDGX agent

Zscaler Inc. is acquiring Symmetry Systems Inc., a venture-backed startup that helps companies monitor their data assets. The companies announced the acquisition today without disclosing the financial

20 May 2026

A Closed-loop, State-centric, Multi-agent Framework for Passenger Load Estimation from Heterogeneous Data Streams

AgentsDGX agent

arXiv:2605.19834v1 Announce Type: cross Abstract: To support operations and passenger-facing services, transit agencies need reliable passenger load trajectories. Currently, load estimates are typical

A novel YOLO26-MoE optimized by an LLM agent for insulator fault detection considering UAV images

AgentsDGX agent

arXiv:2605.19595v1 Announce Type: cross Abstract: The inspection of electrical power line insulators is essential for ensuring grid reliability and preventing failures caused by damaged or degraded in

Agentic Discovery of Cryomicroneedle Formulations

AgentsDGX agent

arXiv:2605.19677v1 Announce Type: new Abstract: Cryomicroneedles offer a route to minimally invasive intradermal delivery of living cells, but their cryogenic formulations must reconcile cell protecti

CMAD: Cooperative Multi-Agent Diffusion via Stochastic Optimal Control

AgentsDGX agent

arXiv:2602.10933v2 Announce Type: replace Abstract: Continuous-time generative models have achieved remarkable success in image restoration and synthesis. However, controlling the composition of multi

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

Model ReleasesDGX agent

arXiv:2605.19484v1 Announce Type: cross Abstract: While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workfl

Google Flow adds agentic brainstorming, more precise editing tools and sharing features

AgentsDGX agent

Google Flow was one of the headline announcements at last year’s I/O event, and it got more attention during this year’s edition today. Flow is a creative studio that’s powered by various artificial i

Governing AI agents at scale with Unity Catalog

IndustryDGX agent

Unity Catalog is Databricks' centralized governance solution that enables organizations to manage and control AI agents at scale by providing features for access control, lineage tracking, and complia

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

SafetyDGX agent

arXiv:2603.11768v2 Announce Type: replace Abstract: Long-term memory has emerged as a foundational component of autonomous Large Language Model (LLM) agents, enabling continuous adaptation, lifelong m

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

Model ReleasesDGX agent

arXiv:2605.19192v1 Announce Type: new Abstract: Multimodal agents use screenshots, documents, and webpages to choose tool calls. When a false visual claim triggers a click, email, extraction, or trans

If Google can’t make AI agents useful, maybe no one can

AgentsDGX agent

For years, tech companies have promised AI will give everyone a capable personal assistant but delivered something more like a clueless intern. Over the past six months, that has started to change, th

Personally incredibly exciting at LangChain Labs to lead a focused, applied research effort into Continual Learning. A big prior is in treat…

AgentsDGX agent

Personally incredibly exciting at LangChain Labs to lead a focused, applied research effort into Continual Learning. A big prior is in treating the long horizon agent optimization problem as one that

POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

Model ReleasesDGX agent

arXiv:2605.19127v1 Announce Type: new Abstract: LLM agents increasingly have access to private user data and act on the user's behalf when interacting with third-party systems. The user defines what m

Riverbed adds agentic AI, session replay and AI observability to Aternity platform

AgentsDGX agent

Digital employee experience company Riverbed Technology LLC today rolled out six additions to its Aternity Experience Platform that the company says move enterprise information technology closer to au

Towards Discovery of Polymers for Insulin Delivery via Physics-Grounded Agentic Workflows

AgentsDGX agent

arXiv:2605.18831v1 Announce Type: cross Abstract: Cold-chain storage limits access to insulin for hundreds of millions of people; a thermally protective patch polymer could help, but the design space

19 May 2026

Agentic Cost-Aware Query Planning with Knowledge Distillation for Big Data Analytics

AgentsDGX agent

arXiv:2605.17831v1 Announce Type: new Abstract: Query optimization in big data analytics remains computationally expensive, particularly for resource-constrained environments where traditional optimiz

Agentmw: Open-source middleware for AI agents — catches mid-run failures,compresses stale context, and grows a reasoning library across runs. Any model, any framework.

Local AiDGX agent

Agentmw is an open-source middleware framework designed to enhance AI agent reliability and efficiency across different models and frameworks. It addresses key operational challenges including mid-run

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

Model ReleasesDGX agent

arXiv:2605.17583v1 Announce Type: new Abstract: While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

ResearchDGX agent

arXiv:2605.17933v1 Announce Type: new Abstract: Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Model ReleasesDGX agent

arXiv:2605.16679v1 Announce Type: cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Model ReleasesDGX agent

arXiv:2511.20857v2 Announce Type: replace-cross Abstract: Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critica

Google targets AI agents and video generation with Gemini 3.5 Flash and Omni

Model ReleasesDGX agent

Google LLC today introduced two new generative artificial intelligence models that push its Gemini family further into AI agents and multimodal creation: Gemini 3.5 Flash, a fast reasoning model desig

Google’s new Gemini Spark is an always-on AI agent for daily digital tasks

Model ReleasesDGX agent

Entering into the artificial intelligence agent era, Google LLC today introduced Gemini Spark, a 24/7 personal AI assistant that can help people navigate their digital lives and do real work on their

Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes

SafetyDGX agent

arXiv:2605.16268v1 Announce Type: cross Abstract: Banks receive millions of reports of fraud, scams, and disputed transactions every year, making it challenging to accurately direct customers to the a

KISS - Knowledge Infrastructure for Scientific Simulation: A Scaffolding for Agentic Earth Science

Model ReleasesDGX agent

arXiv:2605.17856v1 Announce Type: new Abstract: Process-based simulation models encode decades of scientific understanding across the Earth sciences, yet the communities most exposed to climate risk a

LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

Model ReleasesDGX agent

arXiv:2605.18565v1 Announce Type: cross Abstract: Real-world agents operate over long and evolving horizons, where information is repeatedly updated and may interfere across memories, requiring accura

OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents

Model ReleasesDGX agent

arXiv:2506.16042v2 Announce Type: replace Abstract: Generative AI is being leveraged to solve a variety of computer-use tasks involving desktop applications. State-of-the-art systems have focused sole

PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows

Local AiDGX agent

arXiv:2605.18032v1 Announce Type: cross Abstract: Multi-agent LLM workflows -- systems composed of multiple role-specific LLM calls -- often outperform single-prompt baselines, but they remain difficu

R2V Agent: Teaching SLMs When to Ask for Help

Local AiDGX agent

arXiv:2605.16604v1 Announce Type: new Abstract: Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM casc

SE-GA: Memory-Augmented Self-Evolution for GUI Agents

Model ReleasesDGX agent

arXiv:2605.16883v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents often struggle with multi-step tasks due to constrained context windows and static policies that fail t

STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery

AgentsDGX agent

arXiv:2605.17790v1 Announce Type: new Abstract: LLM-based equation discovery offers a promising route to recovering symbolic laws from data, but many systems still rely on generation-centered loops th

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

Model ReleasesDGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

Model ReleasesDGX agent

arXiv:2605.17637v1 Announce Type: new Abstract: Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate tr

18 May 2026

Agent4POI: Agentic Context-Conditioned Affordance Reasoning for Multimodal Point-of-Interest Recommendation

AgentsDGX agent

arXiv:2605.15203v1 Announce Type: cross Abstract: We introduce Agent4POI, the first POI recommendation framework that generates context-conditioned multimodal representations at recommendation time, r

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

SafetyDGX agent

arXiv:2605.15565v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly used to improve the reasoning, coding, and tool-use capabilities of large language models, but agentic RL

Autonomous Intelligent Agents for Natural-Language-Driven Web Execution with Integrated Security Assurance

AgentsDGX agent

arXiv:2605.15281v1 Announce Type: cross Abstract: Modern web test suites rot. A UI refactor breaks locators, a timing change causes race conditions, and within weeks developers abandon the suite entir

FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast

Model ReleasesDGX agent

arXiv:2605.16233v1 Announce Type: new Abstract: Can LLM agents improve decision-making through self-generated memory without gradient updates? We propose FORGE (Failure-Optimized Reflective Graduation

Redis debuts the much-needed memory layer for enterprise AI agents

ApplicationsDGX agent

Artificial intelligence agents have a memory problem and now Redis Inc., the database management startup, is trying to fix that with its new, real-time Context Engine. As the company explains, it’s al

Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation

Model ReleasesDGX agent

arXiv:2605.15669v1 Announce Type: new Abstract: Manufacturable chip layouts must satisfy thousands of geometry-based design rules, and design rule checking (DRC) enforces them by running executable DR

Sigma Computing seals $80M funding round as it pivots toward ‘agentic analytics’

AgentsDGX agent

Cloud-native data analytics startup Sigma Computing Inc. has closed on an 80 million Series E funding round that doubles its valuation to 3 billion, almost a year to the day after its previous Series

SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces

Model ReleasesDGX agent

arXiv:2605.15215v1 Announce Type: new Abstract: Recently, skills have been widely adopted in large language model (LLM)-based agent systems across various domains. In existing frameworks, skills are t

There are a lot of coding and reasoning benchmarks for AI agents, but not a lot for document understanding - which is a prerequisite for all…

Model ReleasesDGX agent

There are a lot of coding and reasoning benchmarks for AI agents, but not a lot for document understanding - which is a prerequisite for all downstream knowledge work. We released ParseBench ~a month

UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization

Local AiDGX agent

arXiv:2510.03161v2 Announce Type: replace-cross Abstract: With the rapid advancements in image generation, synthetic images have become increasingly realistic, posing significant societal risks, such

15 May 2026

A Tutorial on Cognitive Biases in Agentic AI-Driven 6G Autonomous Networks

Model ReleasesDGX agent

arXiv:2510.19973v4 Announce Type: replace-cross Abstract: The path to higher network autonomy in 6G lies beyond the mere optimization of key performance indicators (KPIs), requiring systems that perce

Agentic Recommender System with Hierarchical Belief-State Memory

Model ReleasesDGX agent

arXiv:2605.14401v1 Announce Type: cross Abstract: Memory-augmented LLM agents have advanced personalized recommendation, yet existing approaches universally adopt flat memory representations that conf

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

Model ReleasesDGX agent

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Model ReleasesDGX agent

arXiv:2605.14498v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems t

← Previous
1…107108109110111…300
Next →