AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,920 results
29 Jul 2026

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face

AgentsDGX agent

The AI agent that escaped from OpenAI and hacked developer platform Hugging Face attacked other companies as well, OpenAI revealed on Tuesday. The update substantially widens the scope of an already c

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

SafetyDGX agent

arXiv:2607.24850v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning h

The Future of Agentic AI Depends on Cloud Cost Optimization

AgentsDGX agent

Capitalizing on agentic AI depends on the right infrastructure investments, yet enterprises struggle to reduce cloud costs. Vultr VX1™ Cloud Compute presents a way to free up budget for CPUs and GPUs

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

Model ReleasesDGX agent

arXiv:2607.25765v1 Announce Type: new Abstract: Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs

28 Jul 2026

AlloBench: Measuring Online Tool Allocation Capability in LLM Agents

Model ReleasesDGX agent

arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

Model ReleasesDGX agent

arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated

CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning

SafetyDGX agent

arXiv:2505.18334v2 Announce Type: replace-cross Abstract: Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication is

Knowledge-Centric Agents for Workflow Generation in ComfyUI

AgentsDGX agent

arXiv:2607.15845v2 Announce Type: replace Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular comp

Let AI Agents Translate Networks, Not Reason About Them

Local AiDGX agent

arXiv:2607.22947v1 Announce Type: new Abstract: A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

Model ReleasesDGX agent

arXiv:2607.23870v1 Announce Type: cross Abstract: Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follo

Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View

AgentsDGX agent

arXiv:2607.23029v1 Announce Type: cross Abstract: Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a persistent c

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

SafetyDGX agent

arXiv:2607.24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are tra

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

Model ReleasesDGX agent

arXiv:2508.00288v5 Announce Type: replace-cross Abstract: Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructu

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

SafetyDGX agent

arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whe

27 Jul 2026

Together gives developers an efficient, high-throughput production path for K3’s long, tool-heavy agent workloads, hosted on Together AI’s U…

AgentsDGX agent

Together gives developers an efficient, high-throughput production path for K3’s long, tool-heavy agent workloads, hosted on Together AI’s US-based infrastructure with zero data retention. Start build

Very cool paper from Microsoft. The idea is to train agents on replayed teacher trajectories instead of live environment rollouts. On-policy…

SafetyDGX agent

Very cool paper from Microsoft. The idea is to train agents on replayed teacher trajectories instead of live environment rollouts. On-policy distillation for agentic tasks is expensive because every u

25 Jul 2026

web archive for agent testing

AgentsDGX agent

web archive for agent testing 🔍 Introducing BackSearch. LLMs are increasingly asked to predict the future, but a good backtest requires a snapshot of the internet at a point in time. BackSearch allows

24 Jul 2026

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Model ReleasesDGX agent

arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

SafetyDGX agent

arXiv:2607.21013v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new leve

Frontier Financial Judgement: Can agents tell what might move a stock?

Model ReleasesDGX agent

arXiv:2607.20645v1 Announce Type: cross Abstract: We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents'

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

Model ReleasesDGX agent

arXiv:2607.21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable co

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

Model ReleasesDGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

AgentsDGX agent

arXiv:2511.05385v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agenti

23 Jul 2026

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

Model ReleasesDGX agent

arXiv:2607.19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose

Evaluating AI Agents: A production blueprint with Strands and AgentCore

TutorialsDGX agent

Together, Motorway and AWS built an end-to-end evaluation pipeline that reduced incorrect results from 1 in 8 queries to 1 in 50 and cut issue detection time from few hours to few minutes. The pipelin

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage

SafetyDGX agent

arXiv:2607.19899v1 Announce Type: cross Abstract: Disagreement-triggered escalation can create a structural blind spot in multi-agent arbitration: as base learners improve, they tend to converge, weak

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

AgentsDGX agent

arXiv:2607.19793v1 Announce Type: new Abstract: Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations main

22 Jul 2026

AI Teammates: how monday.com runs production AI agents on Amazon Bedrock

AgentsDGX agent

AI Teammates are agentic AI on Amazon Bedrock, and few engineering organizations run them in production at the scale that monday.com does. Nine in ten Builders use AI coding tools every month, up from

OpenAI said the ‘agent’ escaped a testing environment, gained internet access, stole login credentials and hacked into the start-up Hugging …

AgentsDGX agent

OpenAI said the ‘agent’ escaped a testing environment, gained internet access, stole login credentials and hacked into the start-up Hugging Face by itself — one of the first public examples of a cyber

Simplify AI agent orchestration with Lakebase Postgres

AgentsDGX agent

The article describes a method for creating a scalable AI‑agent orchestrator on Databricks that relies solely on Lakebase Postgres. It explains how the database can handle coordination and task schedu

This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which …

AgentsDGX agent

This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which have been demonstrated in controlled tests for months. Now,

21 Jul 2026

This is a neat feature. I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status…

Model ReleasesDGX agent

This is a neat feature. I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status/2073404610501329247?s=20 But I made it multimodal from the

We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spe…

AgentsDGX agent

We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spent the past 24 hours working closely with the @OpenAI team (

20 Jul 2026

Evolving from legacy BI to agentic AI at Tradeshift with Amazon Quick

AgentsDGX agent

In this post, we describe how Tradeshift deployed Amazon Quick with agentic AI capabilities to replace our legacy BI tool, resulting in query response times up to 30 times faster, a 40 percent reducti

IssueBench is our internal benchmark for evaluating Engine (a continual learning agent in LangSmith) This blog by @nick_bray dives into why …

Model ReleasesDGX agent

**IssueBench** is an internal benchmark created by LangSmith to assess the performance of *Engine*, a continual‑learning agent that scans other agents’ traces to identify, cluster, and fix issues. In

16 Jul 2026

A Self-Evolving Agent for Longitudinal Personal Health Management

Model ReleasesDGX agent

arXiv:2607.13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an ope

AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation

SafetyDGX agent

arXiv:2607.13230v1 Announce Type: new Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interac

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

Model ReleasesDGX agent

arXiv:2607.13465v1 Announce Type: cross Abstract: LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. How

Explaining Reinforcement Learning Agents via Inductive Logic Programming

SafetyDGX agent

arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in saf

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

SafetyDGX agent

arXiv:2604.00830v3 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at

15 Jul 2026

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Model ReleasesDGX agent

arXiv:2607.10350v1 Announce Type: cross Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime la

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

Model ReleasesDGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

Agentic systems for breast cancer treatment recommendations

Model ReleasesDGX agent

arXiv:2607.12051v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

AgentsDGX agent

arXiv:2607.12764v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Re

Rethinking the Evaluation of Harness Evolution for Agents

Model ReleasesDGX agent

arXiv:2607.12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness co

The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning

AgentsDGX agent

arXiv:2607.12177v1 Announce Type: new Abstract: The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

SafetyDGX agent

arXiv:2607.12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable e

14 Jul 2026

Concho AI turns enterprise codebases into a knowledge layer for AI agents

ApplicationsDGX agent

Concho AI today introduced its flagship platform, an artificial intelligence platform that understands software development and application work, providing deep semantic understanding and organization

13 Jul 2026

Devin Fusion is live as an agent preview today in Devin Cloud. Try it out today at https://devin.ai

AgentsDGX agent

Devin Fusion has been released as an agent preview within Devin Cloud, available for users to try today. The announcement was posted by Devin (@devin.ai) at 5:06 PM on July 13, 2026 and has already at

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model…

AgentsDGX agent

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model usage is up, but this could also look like a graph of usage

11 Jul 2026

Community Profiles are live. Proof of work for vibe coders. Your profile, your flex: get an activity graph of your agent usage and checkpoin…

AgentsDGX agent

Community Profiles are live. Proof of work for vibe coders. Your profile, your flex: get an activity graph of your agent usage and checkpoints, plus a Replit Power Ranking for Pro users. Log in, claim

10 Jul 2026

Context Graphs for Proactive Enterprise Agents

Model ReleasesDGX agent

arXiv:2607.07721v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wai

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks

Model ReleasesDGX agent

arXiv:2607.07946v1 Announce Type: cross Abstract: DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks fo

For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same…

AgentsDGX agent

For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper). - Forget everything bel

Malaysia Prime Minister Anwar Ibrahim plans to debut an agentic AI avatar of himself within days, which is meant to help the public navigate government services (Saritha Rai/Bloomberg)

AgentsDGX agent

Saritha Rai / Bloomberg: Malaysia Prime Minister Anwar Ibrahim plans to debut an agentic AI avatar of himself within days, which is meant to help the public navigate government services — Malaysia Pri

MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction

AgentsDGX agent

arXiv:2607.08080v1 Announce Type: new Abstract: Aspect Sentiment Triplet Extraction (ASTE) requires jointly identifying (aspect, opinion, sentiment) triples from a given review sentence. While large l

Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs

SafetyDGX agent

arXiv:2607.08193v1 Announce Type: cross Abstract: Open-ended curricula in Reinforcement Learning (RL) aim to train generally-capable agents by identifying tasks that facilitate learning increasingly c

Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents

Local AiDGX agent

arXiv:2607.08395v1 Announce Type: cross Abstract: Persistent AI agents extend large language models (LLMs) beyond single-turn interaction into long-lived software systems. Unlike traditional chat assi

9 Jul 2026

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Model ReleasesDGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.07178v1 Announce Type: cross Abstract: Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, exis

← Previous
1…7778798081…299
Next →