AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
4 Jun 2026

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

Model ReleasesDGX agent

arXiv:2606.04037v1 Announce Type: new Abstract: Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability bench

Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions

Model ReleasesDGX agent

arXiv:2606.04779v1 Announce Type: new Abstract: Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this

3 Jun 2026

DMF: A Deterministic Memory Framework for Conversational AI Agents

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2606.03463v1 Announce Type: new Abstract: Conversational AI agents require memory systems that are both scalable and semantically coherent across long interaction horizons. Existing approaches r

MemVerse: Multimodal Memory for Lifelong Learning Agents

ResearchDGX agent

arXiv:2512.03627v2 Announce Type: replace Abstract: Despite rapid progress in large-scale language and vision models, AI agents still suffer from a fundamental limitation: they cannot remember. Withou

MUSE: A Unified Agentic Harness for MLLMs

AgentsDGX agent

arXiv:2606.03005v1 Announce Type: cross Abstract: Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze fr

SCOPE: Real-Time Natural Language Camera Agent at the Edge

Model ReleasesDGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

SkillPyramid: A Hierarchical Skill Consolidation Framework for Self-Evolving Agents

ResearchDGX agent

arXiv:2606.03692v1 Announce Type: new Abstract: Recent AI agents can flexibly invoke skills to solve complex tasks, but their long-term improvement is fundamentally constrained by a lack of systematic

Snowflake CoWork brings the agentic enterprise to life for social media marketing

AgentsDGX agent

As AI moves from analytical tool to active participant in daily business operations, solutions like Snowflake CoWork are helping to close the gap between data insight and real-world execution faster t

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

Model ReleasesDGX agent

arXiv:2606.02624v1 Announce Type: cross Abstract: AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier mode…

Model ReleasesDGX agent

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier models on quality and cost by routing selectively to a frontier

2 Jun 2026

A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning

Model ReleasesDGX agent

arXiv:2606.00922v1 Announce Type: cross Abstract: In this work, we propose a prototype machine-to-machine (M2M) knowledge-guided Large Language Model (LLM) framework for automated radiotherapy treatme

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

Model ReleasesDGX agent

arXiv:2512.00986v3 Announce Type: replace Abstract: A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, exi

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

Model ReleasesDGX agent

arXiv:2512.16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks. We study Tools Orchestration Privacy Risk (TOP-R): an age

BranPO: Scalable Contrastive Branch Sampling for Long-Horizon Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2602.03719v2 Announce Type: replace Abstract: Agentic reinforcement learning enables large language models to perform multi-turn planning and tool use, but long-horizon training remains challeng

CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback

SafetyDGX agent

arXiv:2606.01830v1 Announce Type: new Abstract: Recent LLM search agents use reinforcement learning with verifiable rewards (RLVR) to learn search-augmented reasoning from outcome rewards. On hard pro

Digital Twin-Assisted Adaptive Multi-Agent DRL for Intelligent Spectrum and Resource Management in Open-RAN UAV-Enabled 6G Networks

AgentsDGX agent

arXiv:2606.01324v1 Announce Type: cross Abstract: The evolution toward 6G wireless networks envisions a seamlessly intelligent, Open-RAN-enabled architecture where unmanned aerial vehicles (UAVs) play

Dynamic Coordination Strategy Selection for Enterprise Multi-Agent Systems

SafetyDGX agent

arXiv:2606.00804v1 Announce Type: cross Abstract: Enterprise multi-agent systems increasingly expose multiple coordination patterns, but deployments often lack evidence for when to use consensus, deba

Empathic and agentic artificial intelligence in nursing: perspectives on a human-centered framework for cancer care navigation in the United States

AgentsDGX agent

arXiv:2606.00010v1 Announce Type: cross Abstract: For patients experiencing cancer, nurse navigation can ease the burden of complex care by enhancing coordination of health services and patient outcom

From Features to Actions: Explainability in Traditional and Agentic AI Systems

Local AiDGX agent

arXiv:2602.06841v4 Announce Type: replace Abstract: Over the last decade, Explainable AI has primarily focused on interpreting individual model predictions, producing post-hoc explanations that relate

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Model ReleasesDGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

Learning to Retrieve: Dual-Level Long-Term Memory for Text-to-SQL Agents

Model ReleasesDGX agent

arXiv:2606.00547v1 Announce Type: new Abstract: Interactive text-to-SQL agents solve database tasks through multi-turn interactions involving schema exploration, query execution, feedback interpretati

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.02132v1 Announce Type: new Abstract: Agentic reinforcement learning can induce tool abuse, where models overuse external tools even for queries solvable by internal reasoning. Existing appr

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

Model ReleasesDGX agent

arXiv:2512.07436v3 Announce Type: replace Abstract: Recent advances in large reasoning models LRMs have enabled agentic search systems to perform complex multi-step reasoning across multiple sources.

Microsoft releases Web IQ, a search service for AI agents that is powered by Bing, currently used by Copilot, ChatGPT, and other platforms (Barry Schwartz/Search Engine Land)

IndustryDGX agent

Barry Schwartz / Search Engine Land: Microsoft releases Web IQ, a search service for AI agents that is powered by Bing, currently used by Copilot, ChatGPT, and other platforms — Table of Contents — Mi

MOC: Multi-Order Communication in LLM-based Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.02359v1 Announce Type: new Abstract: Despite the remarkable progress of Large Language Model (LLM) based Multi-Agent Systems, most research focuses on optimizing coordination topology while

Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era

HardwareDGX agent

NVIDIA and DigitalOcean are collaborating to develop an open infrastructure stack designed to support autonomous AI agents that operate continuously. The initiative emphasizes open-source principles a

Partial Fairness Awareness: Belief-Guided Strategic Mechanism for Strategic Agents

SafetyDGX agent

arXiv:2606.00826v1 Announce Type: new Abstract: Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models. To addre

Policy and World Modeling Co-Training for Language Agents

SafetyDGX agent

arXiv:2606.02388v1 Announce Type: cross Abstract: Reinforcement learning (RL) improves large language model (LLM) agents by teaching them which actions lead to high rewards, but provides little superv

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

AgentsDGX agent

arXiv:2604.15231v2 Announce Type: replace Abstract: Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (

SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

SafetyDGX agent

arXiv:2606.01991v1 Announce Type: new Abstract: As Large Language Model (LLM) agents increasingly leverage the Model Context Protocol (MCP) to operate in complex environments, the expansion of their a

Self-Evolving Hermes Agents: Enterprise AI That Gets Better With Use | Nemotron Labs https://x.com/i/broadcasts/1pJdRRyneOjKW

Model ReleasesDGX agent

This likely describes a framework or system for deploying AI agents that autonomously improve their performance over time through continuous learning and adaptation in enterprise environments. The sel

Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence

AgentsDGX agent

arXiv:2606.01444v1 Announce Type: new Abstract: Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers a

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

SafetyDGX agent

arXiv:2606.02355v1 Announce Type: new Abstract: Long-horizon LLM agents can benefit from reusable skills, yet existing skill-based methods often rely on external skill generators during training or pe

Site4Drug: Predicting Drug-Binding Target Sites with an AI Agent

AgentsDGX agent

arXiv:2606.01816v1 Announce Type: cross Abstract: Selecting where to intervene on a protein (i.e., choosing a targetable site) is often a more ambiguous and failure-prone bottleneck than selecting wha

SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

Model ReleasesDGX agent

arXiv:2606.02540v1 Announce Type: new Abstract: Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party ski

Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults

Model ReleasesDGX agent

arXiv:2505.19489v2 Announce Type: replace Abstract: The Linux kernel is a critical system, serving as the foundation for numerous systems. Bugs in the Linux kernel can cause serious consequences, affe

The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems

SafetyDGX agent

arXiv:2510.10943v2 Announce Type: replace-cross Abstract: Bias in large language models (LLMs) remains a persistent challenge, often leading to stereotyping and unfair treatment across social groups.

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

SafetyDGX agent

arXiv:2606.00611v1 Announce Type: new Abstract: Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moder

TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control

Local AiDGX agent

arXiv:2604.17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, y

TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement Learning

AgentsDGX agent

arXiv:2508.12551v2 Announce Type: replace-cross Abstract: Linux kernel tuning is essential for optimizing operating system (OS) performance, yet remains challenging due to the complex kernel space, sp

VESTA: Visual Exploration with Statistical Tool Agents

Model ReleasesDGX agent

arXiv:2606.00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems lev

VLAMotor: Test-Guided Enhancement of Vision-Language-Action Models via Agent-BasedData Synthesis

AgentsDGX agent

arXiv:2606.00053v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models follow a data-driven paradigm and are constrained by the coverage of training data, making them prone to failure on

1 Jun 2026

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

Model ReleasesDGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems

HardwareDGX agent

arXiv:2605.30488v1 Announce Type: new Abstract: Collaborative dense SLAM is essential for multi-robot teams to achieve scalable and consistent 3D perception across large-scale outdoor environments. Ex

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

SafetyDGX agent

arXiv:2601.06453v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However,

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

Model ReleasesDGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2605.30712v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong capabilities in reasoning, tool use, and multi-step interaction, but they often solve tasks from scr

🆕Grok Imagine’s Video Agent Moment: Cosmos, xAI, World Models, Generative UI, & the Codex Phase for Video! https://www.latent.space/p/video…

HardwareDGX agent

🆕Grok Imagine’s Video Agent Moment: Cosmos, xAI, World Models, Generative UI, & the Codex Phase for Video! https://www.latent.space/p/video-agents @EthanHe_42, former @xai world model lead and @nvidia

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

AgentsDGX agent

arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected

Safe Equilibrium Policy Optimization for Strategic Agent Policies

Model ReleasesDGX agent

arXiv:2605.30854v1 Announce Type: cross Abstract: Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these age

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories

Model ReleasesDGX agent

arXiv:2605.31308v1 Announce Type: new Abstract: Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introdu

29 May 2026

Agentic AI success helps UiPath swing to a profit, but investors weren’t impressed

AgentsDGX agent

Business automation software company UiPath Inc. delivered mixed results in its latest quarter, posting a solid revenue beat but falling short on earnings — but it did at least manage to return to pro

AIRGuard: Guarding Agent Actions with Runtime Authority Control

SafetyDGX agent

arXiv:2605.28914v1 Announce Type: cross Abstract: Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model C

GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling

AgentsDGX agent

arXiv:2605.28835v1 Announce Type: cross Abstract: Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

Model ReleasesDGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents

Model ReleasesDGX agent

arXiv:2605.29559v1 Announce Type: new Abstract: Mastering terminal environments requires language agents capable of multi-step planning, feedback-grounded execution, and dynamic state adaptation. Howe

No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand

AgentsDGX agent

arXiv:2605.28836v1 Announce Type: cross Abstract: The Plain Writing Act in the United States requires government documents to be accessible in clear and simple language that the general public can eas

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

Model ReleasesDGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

Offline Multi-agent Reinforcement Learning via Sequential Score Decomposition

SafetyDGX agent

arXiv:2505.05968v3 Announce Type: replace Abstract: Offline cooperative multi-agent reinforcement learning (MARL) faces unique challenges due to distributional shifts, particularly stemming from the h

← Previous
1…105106107108109…300
Next →