AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
7 Jul 2026

Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators

Model ReleasesDGX agent

arXiv:2607.04419v1 Announce Type: new Abstract: Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagno

Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator

SafetyDGX agent

arXiv:2509.17255v2 Announce Type: replace-cross Abstract: We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experiments

AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

Model ReleasesDGX agent

arXiv:2604.00392v2 Announce Type: replace-cross Abstract: Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, and

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

SafetyDGX agent

arXiv:2607.05369v1 Announce Type: cross Abstract: For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot progr

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

Model ReleasesDGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

Model ReleasesDGX agent

arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reas

Norm Ai nabs 120M at 1.2B valuation to bring AI agents to the law

ApplicationsDGX agent

Nomos Ai Inc., a company building a platform to embed legal operations into artificial intelligence agents, today announced it has raised 120 million in new funding, bringing the company’s valuation t

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

Model ReleasesDGX agent

arXiv:2607.03968v1 Announce Type: cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine ou

Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets

SafetyDGX agent

arXiv:2607.05179v1 Announce Type: cross Abstract: In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private information

The @better_auth team is joining Vercel to accelerate open source auth for apps and agents. https://vercel.com/blog/vercel-acquires-better-a…

ToolsDGX agent

Vercel has acquired the Better Auth team to accelerate development of open source authentication solutions for applications and AI agents. This acquisition aims to strengthen Vercel's authentication c

The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models

Model ReleasesDGX agent

arXiv:2607.03953v1 Announce Type: cross Abstract: This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the use of struct

this is a great approach, seeing this more @flymy_ai also does this when you build an agent via their api, they'll build a deterministic reu…

Model ReleasesDGX agent

this is a great approach, seeing this more @flymy_ai also does this when you build an agent via their api, they'll build a deterministic reusable workflow, except for where you need models we built th

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

Model ReleasesDGX agent

arXiv:2607.04686v1 Announce Type: cross Abstract: Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never calls a ne

Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation

Model ReleasesDGX agent

arXiv:2603.17673v2 Announce Type: replace-cross Abstract: LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to repro

Transformer-Based Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions

SafetyDGX agent

arXiv:2511.13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations. First,

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.02927v1 Announce Type: cross Abstract: Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (V

6 Jul 2026

Scaling Security Alert Triage With Specialized Agents on Databricks

IndustryDGX agent

This article describes how organizations can use specialized AI agents on the Databricks platform to automate and scale the triage of security alerts, improving the efficiency of security operations t

3 Jul 2026

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Model ReleasesDGX agent

arXiv:2607.01418v1 Announce Type: cross Abstract: Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will ke

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

Model ReleasesDGX agent

arXiv:2607.01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Exist

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Model ReleasesDGX agent

arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern so

Multi-Head Recurrent Memory Agents

ResearchDGX agent

arXiv:2607.01523v1 Announce Type: cross Abstract: Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their sca

The team at @vercel recently released the Eve agent framework, so we built a template that integrates LiteParse with it🦙 The template provi…

Model ReleasesDGX agent

The team at @vercel recently released the Eve agent framework, so we built a template that integrates LiteParse with it🦙 The template provides a set of read-only filesystem tools that let Eve resolve

When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations

SafetyDGX agent

arXiv:2607.01426v1 Announce Type: new Abstract: Autonomous customer-service agents are shifting from conversational interfaces toward operational execution roles: they retrieve firm records, apply ser

2 Jul 2026

AGI Maze as a Benchmark Framework for World-Modeling Agents

Model ReleasesDGX agent

arXiv:2607.00627v1 Announce Type: new Abstract: Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Model ReleasesDGX agent

arXiv:2607.01211v1 Announce Type: cross Abstract: Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real reposi

Finally, Grok’s Speech-to-Text is now live in Grok Build You can now just dictate prompts directly to your coding agents using /voice or Ctr…

IndustryDGX agent

Finally, Grok’s Speech-to-Text is now live in Grok Build You can now just dictate prompts directly to your coding agents using /voice or Ctrl + Space, powered by Grok Voice Just talk naturally for lik

Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows

SafetyDGX agent

arXiv:2607.00269v1 Announce Type: new Abstract: LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale,

Pinecone releases Nexus into public preview to bring business knowledge to AI agents https://ift.tt/4e6Ukbo

ToolsDGX agent

Pinecone has released Nexus, a new product in public preview designed to enable AI agents to access and leverage business knowledge more effectively. Nexus appears to be a solution that integrates wit

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks

Local AiDGX agent

arXiv:2607.00053v1 Announce Type: cross Abstract: Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier m

Urban Deceleration Behavior Modes Under Scene Context: An Early-Kinematic Classifier from Argoverse 2 Multi-Agent Trajectories

AgentsDGX agent

arXiv:2607.00027v1 Announce Type: cross Abstract: Urban deceleration is one of the most empirically studied yet least taxonomically organized behaviors in car-following research. Recent perception-equ

1 Jul 2026

A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases

Model ReleasesDGX agent

arXiv:2606.31041v1 Announce Type: new Abstract: Natural language-to-SQL (NL2SQL) over real-world enterprise databases remains significantly more challenging than on academic benchmarks. Enterprise sch

An Executable Benchmarking Suite for Tool-Using Agents

Model ReleasesDGX agent

arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often con

DA-Studio: An Agentic System for End-to-End Data Analysis

AgentsDGX agent

arXiv:2606.31423v1 Announce Type: cross Abstract: Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system should auto

Enforce consistent code for agents and humans with konsistent

ToolsDGX agent

Vercel introduced Konsistent, a tool designed to enforce consistent code standards across both AI agents and human developers. The solution helps maintain unified coding practices, style guidelines, a

I predict organizations of agents will outperform pure task-based routers on price/performance.

ApplicationsDGX agent

Ethan Mollick predicts that multi-agent organizational structures will deliver better cost-to-performance ratios compared to simple task-based routing systems for AI applications. This suggests that m

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

Model ReleasesDGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

Model ReleasesDGX agent

arXiv:2606.31648v1 Announce Type: new Abstract: We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterp

30 Jun 2026

Agentic Tool Use in Large Language Models

SafetyDGX agent

arXiv:2604.00835v2 Announce Type: replace Abstract: Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for informat

An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations

Model ReleasesDGX agent

arXiv:2606.28467v1 Announce Type: cross Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use. This paper proposes an

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

SafetyDGX agent

arXiv:2606.20470v2 Announce Type: replace-cross Abstract: Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordina

Anthropic launches Claude Sonnet 5, saying it nears Opus 4.8 performance at lower prices and is substantially better than Sonnet 4.6 for agentic work (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic launches Claude Sonnet 5, saying it nears Opus 4.8 performance at lower prices and is substantially better than Sonnet 4.6 for agentic work — Claude Sonnet 5 is built to be the mo

AWS launches forward-deployed engineering team to speed enterprise agentic AI adoption

Model ReleasesDGX agent

Amazon Web Services Inc. said today it’s rolling out a new dedicated organization to bring agentic artificial intelligence systems, built on the same technology, to customers by embedding engineers in

Boundary Degree as a Node-level Feature for Epidemic Scenario Identification in Agent-based Cascade Simulations

AgentsDGX agent

arXiv:2606.29596v1 Announce Type: cross Abstract: Characterizing the scenario underlying an epidemic from its disease cascade is an important task in simulation analytics. We propose boundary degree,

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Model ReleasesDGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Model ReleasesDGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem

SafetyDGX agent

arXiv:2606.29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combin

CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.29476v1 Announce Type: cross Abstract: Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same pol

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

Model ReleasesDGX agent

arXiv:2606.29961v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typi

From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent

Model ReleasesDGX agent

arXiv:2606.30191v1 Announce Type: new Abstract: How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system can detect its

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval

Model ReleasesDGX agent

arXiv:2606.28379v1 Announce Type: cross Abstract: We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structured documents

“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creato…

Model ReleasesDGX agent

“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of h

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes

ResearchDGX agent

arXiv:2606.28900v1 Announce Type: new Abstract: Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they acquire evid

MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation

Model ReleasesDGX agent

arXiv:2606.29914v1 Announce Type: new Abstract: Agent memory systems are increasingly evaluated against RAG and full-context baselines, but reported gains often mix changes in the memory method with c

Powering AI agents: CoreWeave’s validation of Nvidia Vera Rubin signals new chapter for rack-scale computing

HardwareDGX agent

Agentic AI is driving key technology providers to rethink the computing architecture required to run rapidly expanding autonomous systems. In response to this challenge, two leading tech companies hav

Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

Model ReleasesDGX agent

arXiv:2606.18112v3 Announce Type: replace-cross Abstract: Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, becaus

Selective Memory Retention for Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2606.29178v1 Announce Type: new Abstract: When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external memory in froze

Startup OpenMatter wants to make enterprises prove what their AI agents do

ApplicationsDGX agent

OpenMatter Network Inc. today launched a platform it says lets organizations collaborate, run sensitive workloads and deploy artificial intelligence agents across computing environments they do not fu

29 Jun 2026

Build realtime voice agents on AI Gateway

ToolsDGX agent

Vercel announced capabilities for building real-time voice agents using their AI Gateway, enabling developers to create conversational AI applications with voice interaction. The post likely covers te

← Previous
1…114115116117118…300
Next →