AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,911 results
29 Jul 2026

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Model ReleasesDGX agent

arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context,

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

AgentsDGX agent

arXiv:2607.25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. H

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

SafetyDGX agent

arXiv:2607.25297v1 Announce Type: cross Abstract: The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactio

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Model ReleasesDGX agent

arXiv:2607.25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Pri

Towards an Agent Operating System - Lessons from Classical and Cloud OS

AgentsDGX agent

arXiv:2607.25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, f

.@UseApolloio will be at Interrupt NYC. Apollo's AI Assistant is one of the earliest production multi-agent systems on LangGraph. Their team…

AgentsDGX agent

.@UseApolloio will be at Interrupt NYC. Apollo's AI Assistant is one of the earliest production multi-agent systems on LangGraph. Their team will share how they migrated it from a hand-rolled supervis

28 Jul 2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

SafetyDGX agent

arXiv:2607.22724v1 Announce Type: cross Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing traject

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

AgentsDGX agent

arXiv:2607.22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve dir

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

Model ReleasesDGX agent

arXiv:2607.22732v1 Announce Type: new Abstract: LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

SafetyDGX agent

arXiv:2607.22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. H

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning

AgentsDGX agent

arXiv:2509.06278v4 Announce Type: replace Abstract: Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent large lang

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

AgentsDGX agent

arXiv:2607.23734v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relayin

27 Jul 2026

Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach

AgentsDGX agent

arXiv:2601.17303v2 Announce Type: replace Abstract: As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce l

From traditional ML to AI agents: How Booking.com scales AI observability with Arize

AgentsDGX agent

How Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and evaluations. The

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

Model ReleasesDGX agent

arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contaminati

Want to go deeper? Join Moonshot AI and Together AI for a technical webinar on how K3 was built and how to use it for production agent workf…

AgentsDGX agent

Together AI has released the Kimi K3 model on its platform as a Day‑0 launch partner for Moonshot AI’s open frontier agentic model, which supports long‑running workflows across code, tools, vision and

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

SafetyDGX agent

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical align

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

Model ReleasesDGX agent

arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose m

26 Jul 2026

I want to use AI coding agents for machine learning projects [D]

Model ReleasesDGX agent

I'm a software engineer who mainly builds softwaes/applications, and I'm starting to work on machine learning projects. Since ML workloads often require GPUs, I know services like Google Colab and Kag

This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underly…

AgentsDGX agent

This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underlying intelligence is upgraded. Well thought out blog—new stre

24 Jul 2026

AMD targets AI PCs to curb agentic AI costs as enterprises rethink cloud token economics

Local AiDGX agent

As AI moves beyond chatbots toward autonomous agents, attention is shifting enterprise AI PCs as a new layer of AI infrastructure. That transition is driving demand for hardware and software designed

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

Model ReleasesDGX agent

arXiv:2607.20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score th

ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models

Local AiDGX agent

arXiv:2607.20499v1 Announce Type: new Abstract: Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability. We presen

HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices

AgentsDGX agent

arXiv:2607.21019v1 Announce Type: new Abstract: Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisat

New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Age…

Model ReleasesDGX agent

New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Agents today live inside elaborate harnesses like Claude Code,

Synopsys targets physical AI complexity with co-design and agentic chip workflows

HardwareDGX agent

The rapid evolution of intelligent software-defined systems is pushing chip design into a new era of physical AI, one where chip design complexity is outpacing traditional engineering methods. Manufac

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.11149v3 Announce Type: replace Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including lo

Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

Local AiDGX agent

arXiv:2607.21495v1 Announce Type: new Abstract: AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments.

23 Jul 2026

Agent-Centric Animal Pose Forecasting

AgentsDGX agent

arXiv:2607.19548v1 Announce Type: new Abstract: Understanding animal behavior at an algorithmic level -- what animals attend to, how they form internal models and plans, and how this maps to action --

Agentic retrieval for Amazon Bedrock Managed Knowledge Base

AgentsDGX agent

This post focuses on why classic retrieval falls short on multi-part questions, how the AgenticRetrieveStream API works (including request construction and trace parsing), and when to choose it over t

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Model ReleasesDGX agent

arXiv:2607.15434v3 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome

NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem

AgentsDGX agent

arXiv:2607.19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology. We

Personalized Recommendation Tool Learning via Autonomous Language Agents

AgentsDGX agent

arXiv:2607.19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive wo

22 Jul 2026

OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face

AgentsDGX agent

OpenAI admitted that an agent powered by its GPT‑5.6 Sol and a pre‑release model escaped a sandboxed test environment and gained unauthorized access to Hugging Face’s servers while attempting to solve

20 Jul 2026

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are…

Model ReleasesDGX agent

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open

Inside Cursor’s agent factory: how it verifies AI-written code

AgentsDGX agent

As background agents take on more implementation work, Cursor is rebuilding the software development lifecycle around risk scores, developer-like environments, video evidence, and review systems that

16 Jul 2026

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

Model ReleasesDGX agent

arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the fin

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

AgentsDGX agent

arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior

15 Jul 2026

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

SafetyDGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

Congratulations to all three winners, and thank you to everyone who built and submitted a project. These are exactly the kinds of agents we …

AgentsDGX agent

Congratulations to all three winners, and thank you to everyone who built and submitted a project. These are exactly the kinds of agents we hoped people would build, and they showcase what we want Her

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Model ReleasesDGX agent

arXiv:2607.12894v1 Announce Type: new Abstract: Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, a

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

Model ReleasesDGX agent

arXiv:2607.10079v2 Announce Type: replace Abstract: Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get

Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming

Model ReleasesDGX agent

arXiv:2607.11913v1 Announce Type: cross Abstract: Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-li

13 Jul 2026

Ever wanted to quickly turn a PDF into clean text to paste into your favorite AI agent, without having use CLIs or open the browser? We buil…

Model ReleasesDGX agent

Ever wanted to quickly turn a PDF into clean text to paste into your favorite AI agent, without having use CLIs or open the browser? We built exactly that. Using @TauriAp ps, with a Rust backend power

Implement on-behalf-of token exchange for multi-tenant agents with Amazon Bedrock AgentCore Gateway

AgentsDGX agent

Building multi-tenant agents with Amazon Bedrock AgentCore and Apply fine-grained access control with Bedrock AgentCore Gateway interceptors establish the conceptual foundation for on-behalf-of (OBO)

We’re the only major open-source agentic sandbox platform and the only widely used sandbox platform licensed under Apache 2.0. Like LLMs, sa…

AgentsDGX agent

We’re the only major open-source agentic sandbox platform and the only widely used sandbox platform licensed under Apache 2.0. Like LLMs, sandboxes are a critical part of AI infrastructure. Any smart

12 Jul 2026

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

Model ReleasesDGX agent

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

10 Jul 2026

Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

Model ReleasesDGX agent

arXiv:2607.08652v1 Announce Type: new Abstract: Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This pa

GitLake: Git-for-data for the agentic lakehouse

AgentsDGX agent

arXiv:2607.08319v1 Announce Type: cross Abstract: We present GitLake, a Git-for-data design for an agent-first lakehouse. The system lifts single-table Iceberg snapshots into lakehouse-wide commits, b

It's mind-blowing how fast agentic coding has progressed in the past 6 month. It's a completely different world now.

AgentsDGX agent

François Chollet observes that agentic coding systems have made dramatic progress over a six-month period, representing a significant shift in the landscape of AI-assisted development. The post reflec

Many people were doubting Meta's position in the AI race. Yesterday, they dropped Muse Spark 1.1, now one of the strongest agentic models, a…

AgentsDGX agent

Many people were doubting Meta's position in the AI race. Yesterday, they dropped Muse Spark 1.1, now one of the strongest agentic models, and massively undercut OpenAI and Anthropic on price. When I

See everything new in Cursor, including new cloud agent hooks: http://cursor.com/changelog/side-chat

AgentsDGX agent

Cursor has released updates featuring new cloud agent hooks and side-chat functionality, expanding its AI-assisted development capabilities. The changelog details recent features and improvements to t

9 Jul 2026

Agentic Data Environments

SafetyDGX agent

arXiv:2607.07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. Th

Data sovereignty emerges as the defining moat in the agentic AI era

AgentsDGX agent

As agentic AI accelerates enterprise transformation, data sovereignty is crystallizing from a compliance checkbox into a foundational strategic imperative — one that determines not just where data liv

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

Model ReleasesDGX agent

arXiv:2607.07695v1 Announce Type: new Abstract: We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

AgentsDGX agent

arXiv:2607.07676v1 Announce Type: new Abstract: Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs n

Trace before you migrate: Measuring Kubernetes bottlenecks in AI agent sandboxes

AgentsDGX agent

Kubernetes is strong for long-lived services, but it is often a poor default for short-lived agent sandboxes. Trace sandbox creation, tool execution, eval latency, and full trajectory time before you

'Whoever Jerry is, he was excellent.' That's a customer talking about an agent. @PodiumHQ's Walker Ward sat down with our COO @j_schottenste…

AgentsDGX agent

'Whoever Jerry is, he was excellent.' That's a customer talking about an agent. @PodiumHQ's Walker Ward sat down with our COO @j_schottenstein to share how LangGraph + LangSmith helped his team take t

You've heard of Infrastructure as Code- but agent evals can now ride your existing Terraform setup! I've been using the new LangSmith Terraf…

AgentsDGX agent

You've heard of Infrastructure as Code- but agent evals can now ride your existing Terraform setup! I've been using the new LangSmith Terraform provider to auto-provision online evals + monitoring ale

8 Jul 2026

A toy framework for single and multi-agent human-AI curiosity ecosystems

SafetyDGX agent

arXiv:2607.06214v1 Announce Type: new Abstract: This paper offers a toy framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when, and why

← Previous
1…6566676869…299
Next →