AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
9 Jun 2026

Does Persona Make LLMs K-pop Fans? A Pilot Study of LLM-Based Online Concert Audience Agents

SafetyDGX agent

arXiv:2606.07837v1 Announce Type: cross Abstract: A concert is a collective experience, but recorded performance videos are typically watched alone, stripping away the shared audience presence that ma

From one-off prompts to workflows: How to use custom agents in GitHub Copilot CLI

TutorialsDGX agent

Custom agents let GitHub Copilot CLI understand your stack and team workflows, turning one-off terminal prompts into repeatable, reviewable processes. The post From one-off prompts to workflows: How t

GRPO Does Not Close the Multi-Agent Coordination Gap

Model ReleasesDGX agent

arXiv:2606.07845v1 Announce Type: cross Abstract: We measure how well current large language models coordinate as multiple agents sharing a common resource, using the dining philosophers problem as a

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

IRAM-Omega-Q: A Computational Framework for Uncertainty Regulation in Adaptive Agents

ResearchDGX agent

arXiv:2603.16020v2 Announce Type: replace Abstract: Adaptive agents operating under uncertainty must do more than optimize task outputs: they must maintain a workable internal state under noise, pertu

Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2606.09483v1 Announce Type: cross Abstract: Long-term memory for an LLM agent is more than retrieving the right passage at the right time. Current memory systems collapse belief revision, causal

Payoff scaling shapes cooperation in LLM agents across languages

SafetyDGX agent

arXiv:2601.19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users. Whether they coo

SceneConductor: 3D Scene Generation from Single Image with Multi-Agent Orchestration

Model ReleasesDGX agent

arXiv:2606.08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context fro

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

SafetyDGX agent

arXiv:2606.08234v1 Announce Type: new Abstract: LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning

SecureClaw: Clawing Back Control of LLM Agents

SafetyDGX agent

arXiv:2606.09549v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext in

Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control

SafetyDGX agent

arXiv:2606.08405v1 Announce Type: new Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific discovery in physical systems fundamentally requires

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Model ReleasesDGX agent

arXiv:2606.09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However,

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

Model ReleasesDGX agent

arXiv:2606.07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking extern

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08235v2 Announce Type: replace-cross Abstract: Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unin

8 Jun 2026

Audio-Visual World Models: Grounding Multisensory Imagination for Embodied Agents

Model ReleasesDGX agent

arXiv:2512.00883v3 Announce Type: replace-cross Abstract: World models simulate environmental dynamics to enable agents to plan and reason about future states. While existing approaches have primarily

Feasible Action Space Reduction for Quantifying Causal Responsibility in Continuous Spatial Interactions

AgentsDGX agent

arXiv:2505.17739v2 Announce Type: replace-cross Abstract: Understanding the causal influence of one agent on another agent is crucial for safely deploying artificially intelligent systems such as auto

Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.07316v1 Announce Type: cross Abstract: Byzantine collaboration among large-language-model agents requires a finality-control primitive: given delivered stochastic, structured natural-langua

I like when my agents are in tiny windows, so now you can too

ResearchDGX agent

Nous Research has released a feature or tool that allows users to display their AI agents in small, compact window sizes, improving UI flexibility and workspace management. This update addresses user

Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates

SafetyDGX agent

arXiv:2601.18510v2 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights a

M^3Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

Model ReleasesDGX agent

arXiv:2606.07402v1 Announce Type: new Abstract: Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visual

MacArena: Benchmarking Computer Use Agents on an Online macOS Environment

Model ReleasesDGX agent

arXiv:2606.06560v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidl

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents

Local AiDGX agent

arXiv:2606.07027v1 Announce Type: new Abstract: Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but trajectory-le

6 Jun 2026

DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention

Model ReleasesDGX agent

arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

SafetyDGX agent

arXiv:2606.05263v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chai

5 Jun 2026

Augment Code launches Cosmos to bring agentic AI software development to teams

Model ReleasesDGX agent

Augment Code Computing Inc., an artificial intelligence agent platform provider, Thursday announced the launch of Cosmos, a service it says is designed to push beyond the era of individual AI coding a

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

Model ReleasesDGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

Here’s this week’s shipping recap 👇 — Nano Banana 2 & Nano Banana Pro are now GA and available via the Gemini Enterprise Agent Platform, Ge…

Model ReleasesDGX agent

Here’s this week’s shipping recap 👇 — Nano Banana 2 & Nano Banana Pro are now GA and available via the Gemini Enterprise Agent Platform, Gemini API, and in @GoogleAIStudio —Co-Scientist, our new multi

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

Model ReleasesDGX agent

arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substa

VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents

Local AiDGX agent

arXiv:2606.05395v1 Announce Type: new Abstract: Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior. We ar

4 Jun 2026

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

ResearchDGX agent

arXiv:2407.03884v4 Announce Type: replace-cross Abstract: Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and

How Endava is redesigning software delivery around AI agents

TutorialsDGX agent

Endava, a software services company, is leveraging AI agents to fundamentally transform its software delivery processes and workflows. The case study likely demonstrates how the company is implementin

Introducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep…

Model ReleasesDGX agent

Introducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise work

Multi-Agent Next-Best-View Optimization for Risk-Averse Planning

SafetyDGX agent

arXiv:2606.04158v1 Announce Type: new Abstract: Multi-agent Next-Best-View (NBV) selection for safe path planning in uncertain and unknown environments requires informative, safety-aware, and efficien

NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents

Model ReleasesDGX agent

NVIDIA's Nemotron 3 Ultra is a 550B-parameter Mixture-of-Experts model with 55B active parameters, optimized for orchestrating complex, long-running agent workflows by combining frontier reasoning and

PersonaTree: Structured Lifecycle Memory for Person Understanding in LLM Agents

SafetyDGX agent

arXiv:2606.04780v1 Announce Type: new Abstract: Persistent LLM agents require memory representations that make the formation of person understanding explicit across long term interaction. Existing age

RAMPART: Registry-based Agentic Memory with Priority-Aware Runtime Transformation

Model ReleasesDGX agent

arXiv:2606.04628v1 Announce Type: new Abstract: RAMPART is a compile-time memory model and pure in-RAM block registry for LLM-based agents. Context assembly is a programmable runtime operation where c

SaliMory: Orchestrating Cognitive Memory for Conversational Agents

ResearchDGX agent

arXiv:2606.04120v1 Announce Type: cross Abstract: Conversational agents that serve as lifelong companions must maintain persistent memory across all interactions. However, simply expanding context win

The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

SafetyDGX agent

arXiv:2604.07778v2 Announce Type: replace Abstract: Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at le

We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a…

Model ReleasesDGX agent

We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise t

3 Jun 2026

CoreWeave’s Vera Rubin milestone sets stage for theCUBE’s agentic AI coverage

HardwareDGX agent

The agentic AI era is putting new pressure on the infrastructure stack, and CoreWeave Inc.’s latest milestone gives the conversation a sharper edge. This week, the company announced that it has comple

Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing

Model ReleasesDGX agent

arXiv:2606.03618v1 Announce Type: new Abstract: AI-assisted coding agents are bottlenecked by input-token cost. Two pathologies of raw human input drive much of this overhead: tokenization inefficienc

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

SafetyDGX agent

arXiv:2606.03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identifica

Hedge-Bench: Benchmarking Agents on Hard, Realistic Tasks Pertaining to Financial Reasoning

Model ReleasesDGX agent

arXiv:2606.03918v1 Announce Type: new Abstract: AI agents can increasingly handle the mechanical tasks of financial analysis: retrieving documents, calculating formulas, updating spreadsheets. The har

Improve your agent’s tool-calling accuracy with SFT and DPO on Amazon SageMaker AI

AgentsDGX agent

In this post, you learn how to use Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) together to improve the tool-calling accuracy of a small language model (SLM). The example uses

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

Model ReleasesDGX agent

arXiv:2602.05302v3 Announce Type: replace Abstract: We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic

Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation

SafetyDGX agent

arXiv:2606.03963v1 Announce Type: cross Abstract: Deep reinforcement learning has shown strong potential for enabling autonomous robots to learn complex navigational tasks. However, its practical use

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.03467v1 Announce Type: new Abstract: LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

Model ReleasesDGX agent

arXiv:2509.09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In th

VulnAgent-R2: Evidence-Calibrated Multi-Agent Auditing for Repository-Level Vulnerability Detection

Local AiDGX agent

arXiv:2603.13384v2 Announce Type: replace-cross Abstract: Software vulnerabilities often depend on cross-file data flow, build options, framework conventions, and runtime guards, so isolated function

2 Jun 2026

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

Model ReleasesDGX agent

arXiv:2603.14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reason

Agents on a Tree: Pathwise Coordination for Multi-Objective Molecular Optimization

SafetyDGX agent

arXiv:2606.00008v1 Announce Type: new Abstract: Multi-objective molecular optimization requires searching vast chemical spaces under conflicting objectives, where early design decisions strongly const

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

Model ReleasesDGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory

Model ReleasesDGX agent

arXiv:2606.01385v1 Announce Type: cross Abstract: Software architecture design is a critical yet inherently complex and knowledge-intensive phase that requires balancing competing quality attributes a

Build agents you can trust across any framework with open evals and a control standard

TutorialsDGX agent

Learn how Microsoft helps developers build trustworthy AI agents with open evaluations, portable runtime controls, production observability, and security workflows that work across frameworks. The pos

Data agents don't fail at writing SQL. They fail at knowing your business. Schemas show you the columns, but they don't tell you which view …

Model ReleasesDGX agent

Data agents don't fail at writing SQL. They fail at knowing your business. Schemas show you the columns, but they don't tell you which view is canonical for ARR, how often each metric updates, or whic

ExpWeaver: LLM Agents Learn from Experience via Latent RAG

Model ReleasesDGX agent

arXiv:2606.01041v1 Announce Type: new Abstract: Experience learning has achieved promising results in enhancing LLM agent planning and reasoning by integrating past interactions as reusable knowledge.

ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment

Model ReleasesDGX agent

arXiv:2606.00644v1 Announce Type: new Abstract: AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

SafetyDGX agent

arXiv:2606.02373v1 Announce Type: new Abstract: Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

Model ReleasesDGX agent

arXiv:2606.01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly

Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2603.24324v4 Announce Type: replace-cross Abstract: Designing effective auxiliary rewards for cooperative multi-agent systems remains challenging, as misaligned incentives can induce suboptimal

LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models

HardwareDGX agent

arXiv:2606.01838v1 Announce Type: cross Abstract: Agentic language model systems alternate between two structurally distinct step types: structured tool calls (short, deterministic, low perplexity) an

← Previous
1…116117118119120…300
Next →