AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
5 Aug 2026

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Model ReleasesDGX agent

arXiv:2608.03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured file

MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents

AgentsDGX agent

arXiv:2608.03844v1 Announce Type: new Abstract: Memory-augmented LLM agents rely on rich context for long-horizon reasoning and acting, yet their memory modules expose a persistent attack surface for

Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems

SafetyDGX agent

arXiv:2608.03239v1 Announce Type: new Abstract: Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create impli

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Sources: Google is in talks with AI coding agent startup Mechanize on a possible deal, potentially worth $1.5B+, to hire some of its talent and license its tech (Business Insider)

AgentsDGX agent

Business Insider: Sources: Google is in talks with AI coding agent startup Mechanize on a possible deal, potentially worth $1.5B+, to hire some of its talent and license its tech — Google wants its AI

4 Aug 2026

A Forward-Inverse Dynamic Game Framework for Enhanced Multi-Agent Trajectory Planning

SafetyDGX agent

arXiv:2608.01636v1 Announce Type: new Abstract: This paper studies feedback Nash equilibrium (FBNE) seeking for multi-agent trajectory planning in nonlinear dynamical systems with unknown agents' obje

A US appeals court overturns a ruling that had temporarily barred Perplexity from using its agentic shopping tools on Amazon's platform (Blake Brittain/Reuters)

AgentsDGX agent

Blake Brittain / Reuters: A US appeals court overturns a ruling that had temporarily barred Perplexity from using its agentic shopping tools on Amazon's platform — A U.S. appeals court on Tuesday over

Israeli startup Zenity bags $125M in funding to build the security layer for AI agents

AgentsDGX agent

Artificial intelligence security startup Zenity Ltd. said today it has closed on a 125 million Series C round of funding to expand its platform, which secures autonomous agents in production across la

Obsidian Security raises $85M as AI agents create cybersecurity’s next major attack surface

AgentsDGX agent

Obsidian Security Inc. has raised an 85 million Series D funding round at a post-money valuation of 1.1 billion as enterprises increasingly look to secure autonomous artificial intelligence agents acc

Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents

SafetyDGX agent

arXiv:2508.08645v3 Announce Type: replace Abstract: As multimodal large language models advance rapidly, the automation of mobile tasks has become increasingly feasible through the use of mobile-use a

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

Model ReleasesDGX agent

arXiv:2608.00485v1 Announce Type: new Abstract: Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods u

When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems

AgentsDGX agent

arXiv:2608.01085v1 Announce Type: cross Abstract: LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduce

3 Aug 2026

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

AgentsDGX agent

arXiv:2512.05131v2 Announce Type: replace-cross Abstract: Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather

Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO

Model ReleasesDGX agent

arXiv:2607.28679v1 Announce Type: new Abstract: Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial inspection in fa

2 Aug 2026

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and sever…

HardwareDGX agent

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of opt

31 Jul 2026

Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories

Model ReleasesDGX agent

arXiv:2607.27595v1 Announce Type: new Abstract: Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-pa

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Model ReleasesDGX agent

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil

(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations

AgentsDGX agent

arXiv:2607.26201v1 Announce Type: cross Abstract: Security operations centers rely on anomaly detection systems to flag suspicious events. Feature-level explanations for anomaly detectors offer limite

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

AgentsDGX agent

arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal ev

29 Jul 2026

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

SafetyDGX agent

arXiv:2607.24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observa

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Model ReleasesDGX agent

arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context,

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

AgentsDGX agent

arXiv:2607.25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. H

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

SafetyDGX agent

arXiv:2607.25297v1 Announce Type: cross Abstract: The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactio

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Model ReleasesDGX agent

arXiv:2607.25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Pri

Towards an Agent Operating System - Lessons from Classical and Cloud OS

AgentsDGX agent

arXiv:2607.25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, f

.@UseApolloio will be at Interrupt NYC. Apollo's AI Assistant is one of the earliest production multi-agent systems on LangGraph. Their team…

AgentsDGX agent

.@UseApolloio will be at Interrupt NYC. Apollo's AI Assistant is one of the earliest production multi-agent systems on LangGraph. Their team will share how they migrated it from a hand-rolled supervis

28 Jul 2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

SafetyDGX agent

arXiv:2607.22724v1 Announce Type: cross Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing traject

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

AgentsDGX agent

arXiv:2607.22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve dir

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

Model ReleasesDGX agent

arXiv:2607.22732v1 Announce Type: new Abstract: LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

SafetyDGX agent

arXiv:2607.22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. H

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning

AgentsDGX agent

arXiv:2509.06278v4 Announce Type: replace Abstract: Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent large lang

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

AgentsDGX agent

arXiv:2607.23734v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relayin

27 Jul 2026

Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach

AgentsDGX agent

arXiv:2601.17303v2 Announce Type: replace Abstract: As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce l

From traditional ML to AI agents: How Booking.com scales AI observability with Arize

AgentsDGX agent

How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and evaluations. The

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

Model ReleasesDGX agent

arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contaminati

Want to go deeper? Join Moonshot AI and Together AI for a technical webinar on how K3 was built and how to use it for production agent workf…

AgentsDGX agent

Together AI has released the Kimi K3 model on its platform as a Day‑0 launch partner for Moonshot AI’s open frontier agentic model, which supports long‑running workflows across code, tools, vision and

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

SafetyDGX agent

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical align

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

Model ReleasesDGX agent

arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose m

26 Jul 2026

I want to use AI coding agents for machine learning projects [D]

Model ReleasesDGX agent

I'm a software engineer who mainly builds softwaes/applications, and I'm starting to work on machine learning projects. Since ML workloads often require GPUs, I know services like Google Colab and Kag

This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underly…

AgentsDGX agent

This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underlying intelligence is upgraded. Well thought out blog—new stre

24 Jul 2026

AMD targets AI PCs to curb agentic AI costs as enterprises rethink cloud token economics

Local AiDGX agent

As AI moves beyond chatbots toward autonomous agents, attention is shifting enterprise AI PCs as a new layer of AI infrastructure. That transition is driving demand for hardware and software designed

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

Model ReleasesDGX agent

arXiv:2607.20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score th

ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models

Local AiDGX agent

arXiv:2607.20499v1 Announce Type: new Abstract: Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability. We presen

HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices

AgentsDGX agent

arXiv:2607.21019v1 Announce Type: new Abstract: Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisat

New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Age…

Model ReleasesDGX agent

New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Agents today live inside elaborate harnesses like Claude Code,

Synopsys targets physical AI complexity with co-design and agentic chip workflows

HardwareDGX agent

The rapid evolution of intelligent software-defined systems is pushing chip design into a new era of physical AI, one where chip design complexity is outpacing traditional engineering methods. Manufac

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.11149v3 Announce Type: replace Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including lo

Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

Local AiDGX agent

arXiv:2607.21495v1 Announce Type: new Abstract: AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments.

23 Jul 2026

Agent-Centric Animal Pose Forecasting

AgentsDGX agent

arXiv:2607.19548v1 Announce Type: new Abstract: Understanding animal behavior at an algorithmic level -- what animals attend to, how they form internal models and plans, and how this maps to action --

Agentic retrieval for Amazon Bedrock Managed Knowledge Base

AgentsDGX agent

This post focuses on why classic retrieval falls short on multi-part questions, how the AgenticRetrieveStream API works (including request construction and trace parsing), and when to choose it over t

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Model ReleasesDGX agent

arXiv:2607.15434v3 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome

NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem

AgentsDGX agent

arXiv:2607.19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology. We

Personalized Recommendation Tool Learning via Autonomous Language Agents

AgentsDGX agent

arXiv:2607.19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive wo

22 Jul 2026

OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face

AgentsDGX agent

OpenAI admitted that an agent powered by its GPT‑5.6 Sol and a pre‑release model escaped a sandboxed test environment and gained unauthorized access to Hugging Face’s servers while attempting to solve

20 Jul 2026

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are…

Model ReleasesDGX agent

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open

Inside Cursor’s agent factory: how it verifies AI-written code

AgentsDGX agent

As background agents take on more implementation work, Cursor is rebuilding the software development lifecycle around risk scores, developer-like environments, video evidence, and review systems that

16 Jul 2026

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

Model ReleasesDGX agent

arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the fin

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

AgentsDGX agent

arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior

15 Jul 2026

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

SafetyDGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

Congratulations to all three winners, and thank you to everyone who built and submitted a project. These are exactly the kinds of agents we …

AgentsDGX agent

Congratulations to all three winners, and thank you to everyone who built and submitted a project. These are exactly the kinds of agents we hoped people would build, and they showcase what we want Her

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Model ReleasesDGX agent

arXiv:2607.12894v1 Announce Type: new Abstract: Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, a

← Previous
1…6465666768…297
Next →