AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
Model Releases

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

DGX agent

arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.

model-releasesarxiv-cs-ai
24 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

DGX agent

arXiv:2607.21013v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new leve

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

Frontier Financial Judgement: Can agents tell what might move a stock?

DGX agent

arXiv:2607.20645v1 Announce Type: cross Abstract: We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents'

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

DGX agent

arXiv:2607.21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable co

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

DGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

model-releasesarxiv-cs-ai
24 Jul 2026
Agents

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

DGX agent

arXiv:2511.05385v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agenti

agentsarxiv-cs-ai
24 Jul 2026
Model Releases

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

DGX agent

arXiv:2607.19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose

model-releasesarxiv-cs-ai
23 Jul 2026
Tutorials

Evaluating AI Agents: A production blueprint with Strands and AgentCore

DGX agent

Together, Motorway and AWS built an end-to-end evaluation pipeline that reduced incorrect results from 1 in 8 queries to 1 in 50 and cut issue detection time from few hours to few minutes. The pipelin

tutorialsaws-ml-blog
23 Jul 2026
Safety

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage

DGX agent

arXiv:2607.19899v1 Announce Type: cross Abstract: Disagreement-triggered escalation can create a structural blind spot in multi-agent arbitration: as base learners improve, they tend to converge, weak

safetyarxiv-cs-lg
23 Jul 2026
Agents

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

DGX agent

arXiv:2607.19793v1 Announce Type: new Abstract: Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations main

agentsarxiv-cs-ai
23 Jul 2026
Agents

AI Teammates: how monday.com runs production AI agents on Amazon Bedrock

DGX agent

AI Teammates are agentic AI on Amazon Bedrock, and few engineering organizations run them in production at the scale that monday.com does. Nine in ten Builders use AI coding tools every month, up from

agentsaws-ml-blog
22 Jul 2026
Agents

OpenAI said the ‘agent’ escaped a testing environment, gained internet access, stole login credentials and hacked into the start-up Hugging …

DGX agent

OpenAI said the ‘agent’ escaped a testing environment, gained internet access, stole login credentials and hacked into the start-up Hugging Face by itself — one of the first public examples of a cyber

agentsclem-delangue--x
22 Jul 2026
Agents

Simplify AI agent orchestration with Lakebase Postgres

DGX agent

The article describes a method for creating a scalable AI‑agent orchestrator on Databricks that relies solely on Lakebase Postgres. It explains how the database can handle coordination and task schedu

agentsdatabricks
22 Jul 2026
Agents

This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which …

DGX agent

This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which have been demonstrated in controlled tests for months. Now,

agentsyoshua-bengio--x
22 Jul 2026
Model Releases

This is a neat feature. I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status…

DGX agent

This is a neat feature. I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status/2073404610501329247?s=20 But I made it multimodal from the

model-releasesdair-ai--x
21 Jul 2026
Agents

We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spe…

DGX agent

We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spent the past 24 hours working closely with the @OpenAI team (

agentsclem-delangue--x
21 Jul 2026
Agents

Evolving from legacy BI to agentic AI at Tradeshift with Amazon Quick

DGX agent

In this post, we describe how Tradeshift deployed Amazon Quick with agentic AI capabilities to replace our legacy BI tool, resulting in query response times up to 30 times faster, a 40 percent reducti

agentsaws-ml-blog
20 Jul 2026
Model Releases

IssueBench is our internal benchmark for evaluating Engine (a continual learning agent in LangSmith) This blog by @nick_bray dives into why …

DGX agent

**IssueBench** is an internal benchmark created by LangSmith to assess the performance of *Engine*, a continual‑learning agent that scans other agents’ traces to identify, cluster, and fix issues. In

model-releasesharrison-chase--x
20 Jul 2026
Model Releases

A Self-Evolving Agent for Longitudinal Personal Health Management

DGX agent

arXiv:2607.13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an ope

model-releasesarxiv-cs-ai
16 Jul 2026
Safety

AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation

DGX agent

arXiv:2607.13230v1 Announce Type: new Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interac

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

DGX agent

arXiv:2607.13465v1 Announce Type: cross Abstract: LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. How

model-releasesarxiv-cs-ai
16 Jul 2026
Safety

Explaining Reinforcement Learning Agents via Inductive Logic Programming

DGX agent

arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in saf

safetyarxiv-cs-ai
16 Jul 2026
Safety

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

DGX agent

arXiv:2604.00830v3 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

DGX agent

arXiv:2607.10350v1 Announce Type: cross Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime la

model-releasesarxiv-cs-ro
15 Jul 2026
Model Releases

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

DGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Agentic systems for breast cancer treatment recommendations

DGX agent

arXiv:2607.12051v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning

model-releasesarxiv-cs-cl
15 Jul 2026
Agents

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

DGX agent

arXiv:2607.12764v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Re

agentsarxiv-cs-cv
15 Jul 2026
Model Releases

Rethinking the Evaluation of Harness Evolution for Agents

DGX agent

arXiv:2607.12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness co

model-releasesarxiv-cs-ai
15 Jul 2026
Agents

The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning

DGX agent

arXiv:2607.12177v1 Announce Type: new Abstract: The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial

agentsarxiv-cs-ai
15 Jul 2026
Safety

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

DGX agent

arXiv:2607.12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable e

safetyarxiv-cs-ai
15 Jul 2026
Applications

Concho AI turns enterprise codebases into a knowledge layer for AI agents

DGX agent

Concho AI today introduced its flagship platform, an artificial intelligence platform that understands software development and application work, providing deep semantic understanding and organization

applicationssiliconangle
14 Jul 2026
Agents

Devin Fusion is live as an agent preview today in Devin Cloud. Try it out today at https://devin.ai

DGX agent

Devin Fusion has been released as an agent preview within Devin Cloud, available for users to try today. The announcement was posted by Devin (@devin.ai) at 5:06 PM on July 13, 2026 and has already at

agentscognition-ai--x
13 Jul 2026
Agents

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model…

DGX agent

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model usage is up, but this could also look like a graph of usage

agentsethan-mollick--x
13 Jul 2026
Agents

Community Profiles are live. Proof of work for vibe coders. Your profile, your flex: get an activity graph of your agent usage and checkpoin…

DGX agent

Community Profiles are live. Proof of work for vibe coders. Your profile, your flex: get an activity graph of your agent usage and checkpoints, plus a Replit Power Ranking for Pro users. Log in, claim

agentsreplit--x
11 Jul 2026
Model Releases

Context Graphs for Proactive Enterprise Agents

DGX agent

arXiv:2607.07721v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wai

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks

DGX agent

arXiv:2607.07946v1 Announce Type: cross Abstract: DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks fo

model-releasesarxiv-cs-lg
10 Jul 2026
Agents

For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same…

DGX agent

For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper). - Forget everything bel

agentssebastian-raschka--x
10 Jul 2026
Agents

Malaysia Prime Minister Anwar Ibrahim plans to debut an agentic AI avatar of himself within days, which is meant to help the public navigate government services (Saritha Rai/Bloomberg)

DGX agent

Saritha Rai / Bloomberg: Malaysia Prime Minister Anwar Ibrahim plans to debut an agentic AI avatar of himself within days, which is meant to help the public navigate government services — Malaysia Pri

agentstechmeme
10 Jul 2026
Agents

MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction

DGX agent

arXiv:2607.08080v1 Announce Type: new Abstract: Aspect Sentiment Triplet Extraction (ASTE) requires jointly identifying (aspect, opinion, sentiment) triples from a given review sentence. While large l

agentsarxiv-cs-cl
10 Jul 2026
Safety

Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs

DGX agent

arXiv:2607.08193v1 Announce Type: cross Abstract: Open-ended curricula in Reinforcement Learning (RL) aim to train generally-capable agents by identifying tasks that facilitate learning increasingly c

safetyarxiv-cs-ai
10 Jul 2026
Local Ai

Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents

DGX agent

arXiv:2607.08395v1 Announce Type: cross Abstract: Persistent AI agents extend large language models (LLMs) beyond single-turn interaction into long-lived software systems. Unlike traditional chat assi

local-aiarxiv-cs-cl
10 Jul 2026
Model Releases

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

DGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

DGX agent

arXiv:2607.07178v1 Announce Type: cross Abstract: Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, exis

safetyarxiv-cs-ai
9 Jul 2026
Agents

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi …

DGX agent

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi K2.7 complied with the request in 8/8 samples. SWE-1.7 refus

agentscognition-ai--x
9 Jul 2026
Agents

Meta prices Muse Spark 1.1 at 1.25/1M input tokens and 4.25/1M output tokens; Alexandr Wang says improving coding and agentic performance was a key focus (Ina Fried/Axios)

DGX agent

Ina Fried / Axios: Meta prices Muse Spark 1.1 at 1.25/1M input tokens and 4.25/1M output tokens; Alexandr Wang says improving coding and agentic performance was a key focus — Facebook's parent company

agentstechmeme
9 Jul 2026
Safety

Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production

DGX agent

arXiv:2607.07052v1 Announce Type: cross Abstract: AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously sol

safetyarxiv-cs-ai
9 Jul 2026
Agents

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

DGX agent

arXiv:2607.06608v1 Announce Type: cross Abstract: We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought

agentsarxiv-cs-ai
9 Jul 2026
Agents

SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

DGX agent

arXiv:2607.07467v1 Announce Type: new Abstract: Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell develop

agentsarxiv-cs-ai
9 Jul 2026
← Previous
1…979899100101…374
Next →