AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
7 Aug 2026

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

SafetyDGX agent

arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible cons

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

Model ReleasesDGX agent

arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local look

From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks

Local AiDGX agent

arXiv:2608.06227v1 Announce Type: cross Abstract: Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.02721v3 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

AgentsDGX agent

arXiv:2608.06161v1 Announce Type: new Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptu

Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

Model ReleasesDGX agent

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs d

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

Model ReleasesDGX agent

arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning error

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

Model ReleasesDGX agent

arXiv:2608.06065v1 Announce Type: new Abstract: GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs:

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

Model ReleasesDGX agent

arXiv:2608.06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Crit

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

AgentsDGX agent

arXiv:2608.05207v1 Announce Type: new Abstract: Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discov

6 Aug 2026

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

Model ReleasesDGX agent

arXiv:2608.04682v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific b

AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering

AgentsDGX agent

arXiv:2411.04440v1 Announce Type: cross Abstract: Protein engineering is important for biomedical applications, but conventional approaches are often inefficient and resource-intensive. While deep lea

Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework

AgentsDGX agent

arXiv:2608.04366v1 Announce Type: cross Abstract: While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabiliti

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

Model ReleasesDGX agent

arXiv:2608.04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompt

General Availability of Pinecone Nexus Proves Knowledge Drives Real Outcomes for Agentic AI

Model ReleasesDGX agent

Pinecone announced the general availability of Pinecone Nexus, a knowledge engine that converts an enterprise’s proprietary data into governed, agent‑ready knowledge delivered through a single query c

Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings for CLEF JOKER 2025 Task 2

AgentsDGX agent

arXiv:2507.06506v2 Announce Type: replace-cross Abstract: Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine tra

5 Aug 2026

CUADebug: Diagnosing and Repairing Computer-Use Agent Failures

Model ReleasesDGX agent

arXiv:2608.02643v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedback, yet thei

Cura 1T: Specialized Model for Agentic Healthcare

Model ReleasesDGX agent

arXiv:2607.15314v2 Announce Type: replace Abstract: Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR)

DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

Model ReleasesDGX agent

arXiv:2608.03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to

From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities

Model ReleasesDGX agent

arXiv:2608.03585v1 Announce Type: new Abstract: Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interp

LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension

AgentsDGX agent

arXiv:2608.02915v1 Announce Type: cross Abstract: Domain-specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC-V ecosystem to accelerate emerging workloads, but implem

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

Model ReleasesDGX agent

arXiv:2608.03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goe

Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory

AgentsDGX agent

arXiv:2608.03420v1 Announce Type: new Abstract: Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well under

Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering

Local AiDGX agent

arXiv:2608.01463v2 Announce Type: replace Abstract: Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Localized Mul

Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

Model ReleasesDGX agent

arXiv:2608.02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts

4 Aug 2026

Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

Model ReleasesDGX agent

arXiv:2608.02391v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution st

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

AgentsDGX agent

arXiv:2608.01827v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to

Ethyca launches Astralis to govern enterprise AI agents in real time

Model ReleasesDGX agent

Data privacy engineering company Ethyca Inc. today launched Astralis, a platform that governs how enterprise artificial intelligence models and agents use company data in real time. The company is pit

Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

SafetyDGX agent

arXiv:2608.02018v1 Announce Type: new Abstract: Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to in

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Model ReleasesDGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

Real-Time Detection and Repair of LLM Agent Failures

Model ReleasesDGX agent

arXiv:2608.02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the stan

Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents

AgentsDGX agent

arXiv:2608.01285v1 Announce Type: new Abstract: The continued development of LLMs toward persistent and adaptive intelligence increasingly requires long-term memory mechanisms that preserve and reuse

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

Model ReleasesDGX agent

arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to thei

3 Aug 2026

HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

AgentsDGX agent

arXiv:2607.29600v1 Announce Type: new Abstract: Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has eme

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

AgentsDGX agent

arXiv:2607.28979v1 Announce Type: new Abstract: Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

HardwareDGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

2 Aug 2026

The $0.01 meeting assistant Records meetings on his phone, sends the audio to Telegram, and his locally-hosted AI agent transcribes, identif…

Local AiDGX agent

The $0.01 meeting assistant Records meetings on his phone, sends the audio to Telegram, and his locally-hosted AI agent transcribes, identifies speakers, extracts action items, and files tasks -- for

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed door…

HardwareDGX agent

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed doors in just a few organizations doesnt work. What worked in th

31 Jul 2026

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

HardwareDGX agent

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a co

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.28076v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only co

30 Jul 2026

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

AgentsDGX agent

arXiv:2607.27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementa

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some inter…

Model ReleasesDGX agent

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively usin

29 Jul 2026

SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation

AgentsDGX agent

arXiv:2607.24802v1 Announce Type: cross Abstract: This paper presents our system for Task 3 of the CLEF 2026 CheckThat! Lab, which focuses on generating full fact-checking articles from claims, veraci

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

Model ReleasesDGX agent

arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), ho

28 Jul 2026

Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence

AgentsDGX agent

arXiv:2607.22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

AgentsDGX agent

arXiv:2607.23693v1 Announce Type: new Abstract: Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and epis

GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks

AgentsDGX agent

arXiv:2607.23792v1 Announce Type: cross Abstract: Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Model ReleasesDGX agent

arXiv:2607.24368v1 Announce Type: new Abstract: Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption

OpenAI says the rogue AI that breached Hugging Face used exposed credentials from 'four accounts' tied to four 'publicly available' third-party services (Wired)

AgentsDGX agent

Wired: OpenAI says the rogue AI that breached Hugging Face used exposed credentials from “four accounts” tied to four “publicly available” third-party services — In a new disclosure, OpenAI says its a

Perplexity’s Personal Computer turns Windows PCs into AI agents

Local AiDGX agent

Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexit

Quoting Akshat Bubna

AgentsDGX agent

We’re aware a Modal customer published an unauthenticated endpoint that allowed ​anyone on the internet to use ​their ⁠sandboxes for code execution. This was used by the rogue agent. Modal’s ⁠platform

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

Model ReleasesDGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

Stress-testing large language model agents in a robotic chemistry laboratory

AgentsDGX agent

arXiv:2607.23045v1 Announce Type: new Abstract: AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. He

27 Jul 2026

Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%…

Model ReleasesDGX agent

Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #1 spot across 5 signals (see below). Kimi

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

AgentsDGX agent

arXiv:2607.22511v1 Announce Type: cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approa

Nexus connects to the agentic harnesses your teams already use, whether that’s Claude Code, Codex, OpenCode, or your own custom tooling. It …

Model ReleasesDGX agent

Nexus connects to the agentic harnesses your teams already use, whether that’s Claude Code, Codex, OpenCode, or your own custom tooling. It gives you: → Intelligent routing, automatically matching eac

NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding

Model ReleasesDGX agent

NVIDIA’s Nemotron 3 Ultra, when paired with the ACE‑RTL agent, achieves a 97.1 % average pass rate on the CVDP benchmark across nine RTL task categories—surpassing GLM 5.2 and Kimi K2.6 while using up

The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games

SafetyDGX agent

arXiv:2603.05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization.

The LangChain podcast where @hwchase17 interviews agent builders is full of alpha. Recent one with @EnoReyes was the best one yet. Sharing m…

Model ReleasesDGX agent

The LangChain podcast where @hwchase17 interviews agent builders is full of alpha. Recent one with @EnoReyes was the best one yet. Sharing my unstructured notes: Eno keeps bringing back some core conc

24 Jul 2026

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

HardwareDGX agent

arXiv:2607.20527v1 Announce Type: new Abstract: Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already

← Previous
1…101102103104105…300
Next →