AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
24 Jun 2026

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

SafetyDGX agent

arXiv:2606.24515v1 Announce Type: new Abstract: Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement lear

Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception

Local AiDGX agent

arXiv:2603.20750v2 Announce Type: replace Abstract: Social actors do not observe a common social world: each individual forms judgments from a partial and potentially distorted view of the surrounding

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent

AgentsDGX agent

arXiv:2606.24094v1 Announce Type: new Abstract: Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VisCritic: Visual State Comparison as Process Reward for GUI Agents

Model ReleasesDGX agent

arXiv:2606.24525v1 Announce Type: new Abstract: GUI agents powered by vision-language models show strong potential for automating digital tasks, yet frequently fail in long-horizon scenarios due to th

23 Jun 2026

CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents

Model ReleasesDGX agent

arXiv:2606.22000v1 Announce Type: cross Abstract: We introduce CFAgentBench, a reproducible, self-hostable environment and benchmark for autonomous construction-finance agents: a CFO/controller-class

Democratizing and accelerating AI-driven pathology research through agentic intelligence

AgentsDGX agent

arXiv:2606.20677v1 Announce Type: cross Abstract: Computational pathology has advanced rapidly with the emergence of foundation models, yet widespread adoption remains limited by substantial technical

From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents

Model ReleasesDGX agent

arXiv:2606.20661v1 Announce Type: cross Abstract: The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize ex

How Telcos Build Autonomous Networks with Agentic AI

HardwareDGX agent

Telecommunication operators deploy autonomous agents built on telecom-domain models and running inside secure execution runtimes connected to tools, digital twins, and shared skills to understand netw

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

Local AiDGX agent

arXiv:2606.20717v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inheren

Momentic raises the bar for software testing with agentic quality platform

AgentsDGX agent

Artificial intelligence-powered software testing and quality assurance platform Momentic Inc. today announced a major update to its service focused on verification in the AI coding era, enabling teams

Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning

SafetyDGX agent

arXiv:2601.20209v2 Announce Type: replace Abstract: Reinforcement learning has empowered large language models to act as intelligent agents, yet training them for long-horizon tasks remains challengin

Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning

AgentsDGX agent

arXiv:2511.14445v2 Announce Type: replace-cross Abstract: We present Tell Me, a mental well-being system that leverages advances in large language models to provide accessible, context-aware support f

22 Jun 2026

GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark GLM-5.2 from @Zai_org scores 1524…

Model ReleasesDGX agent

GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark GLM-5.2 from @Zai_org scores 1524 Elo on GDPval-AA, which measures performance on real-world,

11 Jun 2026

Grounding Computer Use Agents on Human Demonstrations

Model ReleasesDGX agent

arXiv:2511.07332v2 Announce Type: replace-cross Abstract: Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen element

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

Model ReleasesDGX agent

arXiv:2602.22638v2 Announce Type: replace Abstract: Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through na

SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation

SafetyDGX agent

arXiv:2511.05203v3 Announce Type: replace Abstract: Today's autonomous agents, largely driven by foundation models (FMs), can understand natural language instructions and solve long-horizon tasks with

10 Jun 2026

Dmsh: A Multi-Agent Reinforcement Learning Framework for All-Quad Mesh Generation

AgentsDGX agent

arXiv:2606.10601v1 Announce Type: cross Abstract: Generating high-quality meshes for arbitrary geometries remains a fundamental bottleneck in computational engineering, often demanding heuristic tunin

Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate

SafetyDGX agent

arXiv:2606.10307v1 Announce Type: new Abstract: Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whethe

Fact-Augmented Lookahead Planning for LLM Agents

Model ReleasesDGX agent

arXiv:2506.09171v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly capable, but LLM agents still struggle to plan effectively in interactive, partially observable,

Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trace Diagnostics, and a Partial Fix

Model ReleasesDGX agent

arXiv:2606.09884v1 Announce Type: cross Abstract: We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation betw

FinOps AI goes beyond token economics as agentic costs emerge

AgentsDGX agent

As FinOps AI strategies continue to emerge,the familiar cloud cost management approach is breaking down — and organizations that fail to adapt risk runaway spending on workloads they barely understand

GUI-AC: Enhancing Continual Learning in GUI Agents

SafetyDGX agent

arXiv:2606.10522v1 Announce Type: new Abstract: Graphical User Interfaces (GUIs) serve as the dominant medium for human-computer interaction, yet building GUI agents that generalize across the vast di

MIRAGE: A Polarity-Flipping Encoding Subspace in LLM Agents

Model ReleasesDGX agent

arXiv:2606.10304v1 Announce Type: new Abstract: When LLM agents are coerced into covertly encoding sensitive data (Base64, ROT13, acrostic, synonym chains, and beyond), the resulting outputs evade out

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans

Model ReleasesDGX agent

arXiv:2606.09844v1 Announce Type: cross Abstract: Large Language Models (LLMs) alter their privacy behavior based on the perceived identity of their interlocutor. While safety mechanisms typically pre

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. H

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

SafetyDGX agent

arXiv:2606.09421v2 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, valida

9 Jun 2026

Agentic Neuro-Symbolic Planning and Commissioning for Human-in-the-Loop Industrial Robotics with Digital Twins

AgentsDGX agent

arXiv:2606.08214v1 Announce Type: new Abstract: Flexible robotic automation requires systems that interpret operator intent, verify physical feasibility, and recover from execution failures across bot

Bidirectional Semantic Complementary Tool Retrieval for Remote Sensing Agents

Model ReleasesDGX agent

arXiv:2606.07538v1 Announce Type: cross Abstract: Large language model (LLM)-based agents provide a novel paradigm for the automated processing of remote sensing(RS) data. Their success in complex RS

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2606.08151v1 Announce Type: new Abstract: Tool-using LLM agents often fail not because relevant text is absent, but because decisive evidence is not selected, compressed, or surfaced at action t

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination

Model ReleasesDGX agent

arXiv:2606.08068v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems often fail to reliably outperform a single strong model equipped with best-of-N sampling. We argue that a

Earlytrade raises $10M to bring agentic AI to construction payments

AgentsDGX agent

Earlytrade Pty. Ltd., a company solving payments flow for contractors in the construction industry, today announced it has raised about 10 million in new funding. Today’s capital infusion brings the t

How to detect credential theft in AI agent harness traces

AgentsDGX agent

In May 2026, a malicious version of a popular VS Code extension spent 18 minutes in the marketplace before anyone caught it. In that time it ran on roughly 6,000... The post How to detect credential t

iOSWorld: A Benchmark for Personally Intelligent Phone Agents

Model ReleasesDGX agent

arXiv:2606.09764v1 Announce Type: new Abstract: A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device,

PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback

Local AiDGX agent

arXiv:2606.08688v1 Announce Type: cross Abstract: Achieving fully automated, physically plausible 3D motion synthesis is a core objective in graphics and generative AI. However, configuring complex en

Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2606.09610v1 Announce Type: cross Abstract: Cooperative object transportation is essential in numerous domains, including industrial to domestic services. A popular transportation strategy is to

'So There's a Catch-22 Here': How Early Adopters Who Build Multi-Agent LLM Systems Conceptualize Transparency

SafetyDGX agent

arXiv:2606.08323v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) systems are rapidly emerging, yet transparency, a cornerstone of responsible AI, remains under-defined in these

Structuring agentic AI for HPC code modernization

AgentsDGX agent

arXiv:2606.08710v1 Announce Type: cross Abstract: Modernization of legacy scientific codes is often necessary to keep up with the ever-evolving changes in the compute resource ecosystem. Parallelizati

Tiger Data launches PostgreSQL extension designed for AI agents

Model ReleasesDGX agent

Tiger Data today introduced a managed PostgreSQL database service designed specifically for AI agents, saying conventional database architectures are poorly suited to a future in which software is inc

Zscaler launches AI Broker and Endpoint AI Security for agents

Model ReleasesDGX agent

Zscaler Inc. today unveiled a set of products designed to secure autonomous artificial intelligence agents, with the cybersecurity company claiming it has built the industry’s first complete zero-trus

8 Jun 2026

CAF-Gen: A Multi-Agent System for Enriching Argumentation Structures

SafetyDGX agent

arXiv:2606.06646v1 Announce Type: cross Abstract: Formalizing complex reasoning from natural text is one of the central challenges in computational linguistics. It requires systems to understand not j

EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks

SafetyDGX agent

arXiv:2505.14289v2 Announce Type: replace Abstract: Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) are increasingly deployed yet vulnerable to Environmental

Evaluate your Amazon Nova Sonic voice agent at scale, no microphone required

AgentsDGX agent

In this post, we walk you through the Nova Sonic Test Harness, an open source framework that we built to solve both problems. It serves as a rapid iteration tool for tuning system prompts and tool con

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Model ReleasesDGX agent

arXiv:2606.06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recen

Small Language Model Agents Enable Efficient and High-Quality Knowledge Mining

AgentsDGX agent

arXiv:2510.01427v3 Announce Type: replace Abstract: At the core of Deep Research is knowledge mining, the task of extracting structured information from massive unstructured text in response to user i

Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments

Model ReleasesDGX agent

arXiv:2606.06960v1 Announce Type: new Abstract: Experience-based self-evolution is crucial for LLM agents, but existing benchmarks often assume explicit goals, stable task patterns, and clear feedback

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a l…

Model ReleasesDGX agent

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a loop until said goal is complete this is similar to /goal in

6 Jun 2026

Agentic Molecular Recovery via Molecule-Aware Exploration

AgentsDGX agent

arXiv:2606.05847v1 Announce Type: new Abstract: Text-guided molecular generation with LLMs often yields invalid SMILES. We argue that invalid drafts should be addressed through a shift from validity-o

5 Jun 2026

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

Model ReleasesDGX agent

arXiv:2606.05622v1 Announce Type: new Abstract: Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are pro

Beyond tokens: a unified framework for latent communication in LLM-based multi-agent systems

SafetyDGX agent

arXiv:2606.05711v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) have become a prevailing paradigm for tackling complex reasoning, planning, and tool-use tasks

Coding with 'Enemy': Can Human Developers Detect AI Agent Sabotage?

Model ReleasesDGX agent

arXiv:2606.05647v1 Announce Type: cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to cod

EpiEvolve: Self-Evolving Agents for Streaming Pandemic Forecasting under Regime Shifts

AgentsDGX agent

arXiv:2606.05513v1 Announce Type: cross Abstract: Epidemic LLM forecasters are usually trained and evaluated as static supervised models, whereas operational pandemic forecasting is a streaming proces

Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents

Local AiDGX agent

arXiv:2606.05828v1 Announce Type: cross Abstract: As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerg

4 Jun 2026

Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

Model ReleasesDGX agent

arXiv:2606.04701v1 Announce Type: cross Abstract: GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications viol

Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs

AgentsDGX agent

arXiv:2606.04360v1 Announce Type: new Abstract: Symbolic regression (SR) discovers compact mathematical expressions from data, yet recent LLM-based evolutionary methods remain sample-inefficient becau

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-l…

Model ReleasesDGX agent

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now b

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

Model ReleasesDGX agent

arXiv:2603.03205v2 Announce Type: replace Abstract: Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon ac

Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval

Model ReleasesDGX agent

arXiv:2606.04391v1 Announce Type: new Abstract: Language agents increasingly rely on reusable skills to improve multi-step web automation across related tasks. A growing line of work studies online sk

SMADE-IE: Sparse Multi-Agent Framework with Evidence-Driven Debate for Zero-Shot Information Extraction

Model ReleasesDGX agent

arXiv:2606.04691v1 Announce Type: new Abstract: Zero-shot information extraction (IE) with large language models (LLMs) has attracted increasing attention due to its flexibility in adapting to new sch

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

SafetyDGX agent

arXiv:2606.04321v1 Announce Type: new Abstract: Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agen

← Previous
1…104105106107108…300
Next →