AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

DGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Real-Time Detection and Repair of LLM Agent Failures

DGX agent

arXiv:2608.02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the stan

model-releasesarxiv-cs-lg
4 Aug 2026
Agents

Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents

DGX agent

arXiv:2608.01285v1 Announce Type: new Abstract: The continued development of LLMs toward persistent and adaptive intelligence increasingly requires long-term memory mechanisms that preserve and reuse

agentsarxiv-cs-lg
4 Aug 2026
Model Releases

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

DGX agent

arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to thei

model-releasesarxiv-cs-lg
4 Aug 2026
Agents

HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

DGX agent

arXiv:2607.29600v1 Announce Type: new Abstract: Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has eme

agentsarxiv-cs-ro
3 Aug 2026
Agents

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

DGX agent

arXiv:2607.28979v1 Announce Type: new Abstract: Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their

agentsarxiv-cs-cl
3 Aug 2026
Hardware

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

DGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

hardwarearxiv-cs-cl
3 Aug 2026
Hardware

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

DGX agent

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a co

hardwarearxiv-cs-ai
31 Jul 2026
Safety

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

DGX agent

arXiv:2607.28076v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only co

safetyarxiv-cs-lg
31 Jul 2026
Agents

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

DGX agent

arXiv:2607.27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementa

agentsarxiv-cs-cl
30 Jul 2026
Agents

SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation

DGX agent

arXiv:2607.24802v1 Announce Type: cross Abstract: This paper presents our system for Task 3 of the CLEF 2026 CheckThat! Lab, which focuses on generating full fact-checking articles from claims, veraci

agentsarxiv-cs-cl
29 Jul 2026
Model Releases

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

DGX agent

arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), ho

model-releasesarxiv-cs-ai
29 Jul 2026
Agents

Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence

DGX agent

arXiv:2607.22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and

agentsarxiv-cs-ai
28 Jul 2026
Agents

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

DGX agent

arXiv:2607.23693v1 Announce Type: new Abstract: Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and epis

agentsarxiv-cs-ai
28 Jul 2026
Agents

GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks

DGX agent

arXiv:2607.23792v1 Announce Type: cross Abstract: Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion

agentsarxiv-cs-lg
28 Jul 2026
Model Releases

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

DGX agent

arXiv:2607.24368v1 Announce Type: new Abstract: Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

DGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

model-releasesarxiv-cs-ai
28 Jul 2026
Agents

Stress-testing large language model agents in a robotic chemistry laboratory

DGX agent

arXiv:2607.23045v1 Announce Type: new Abstract: AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. He

agentsarxiv-cs-ai
28 Jul 2026
Agents

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

DGX agent

arXiv:2607.22511v1 Announce Type: cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approa

agentsarxiv-cs-lg
27 Jul 2026
Safety

The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games

DGX agent

arXiv:2603.05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization.

safetyarxiv-cs-lg
27 Jul 2026
Hardware

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

DGX agent

arXiv:2607.20527v1 Announce Type: new Abstract: Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already

hardwarearxiv-cs-ai
24 Jul 2026
Safety

Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

DGX agent

arXiv:2607.20708v1 Announce Type: new Abstract: A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that Phi_r grow

safetyarxiv-cs-lg
24 Jul 2026
Agents

Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

DGX agent

arXiv:2607.20431v1 Announce Type: new Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet conv

agentsarxiv-cs-cl
24 Jul 2026
Model Releases

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

DGX agent

arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Te

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

DGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

DGX agent

arXiv:2607.14573v3 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows,

model-releasesarxiv-cs-ai
23 Jul 2026
Agents

Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving

DGX agent

arXiv:2607.19774v1 Announce Type: new Abstract: Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collab

agentsarxiv-cs-ro
23 Jul 2026
Model Releases

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

DGX agent

arXiv:2607.19038v1 Announce Type: new Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

DGX agent

arXiv:2607.14186v4 Announce Type: replace-cross Abstract: Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

DGX agent

arXiv:2607.19351v1 Announce Type: new Abstract: LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions throu

safetyarxiv-cs-ai
23 Jul 2026
Agents

Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation

DGX agent

arXiv:2607.19767v1 Announce Type: new Abstract: A rich and recognizable component library is the cornerstone of printed circuit board (PCB) design and generation. Traditionally, engineers manually cre

agentsarxiv-cs-ai
23 Jul 2026
Safety

The Ethics of Autonomous AI Agents for Offensive Security

DGX agent

arXiv:2607.20255v1 Announce Type: cross Abstract: LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and o

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

DGX agent

arXiv:2607.13792v1 Announce Type: new Abstract: Most daily activities are inherently procedural. However, existing evaluations for egocentric video understanding seldom address procedural understandin

model-releasesarxiv-cs-cv
16 Jul 2026
Agents

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

DGX agent

arXiv:2607.13679v1 Announce Type: new Abstract: AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen o

agentsarxiv-cs-ai
16 Jul 2026
Local Ai

Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition

DGX agent

arXiv:2607.12911v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used for dietary assessment from meal images, where retrieval-augmented grounding was shown to

local-aiarxiv-cs-cv
15 Jul 2026
Agents

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

DGX agent

arXiv:2607.11019v2 Announce Type: replace Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineerin

agentsarxiv-cs-ai
15 Jul 2026
Agents

Unveiling Complex Collective Behaviors from Simple Rewards

DGX agent

arXiv:2607.12861v1 Announce Type: cross Abstract: Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic an

agentsarxiv-cs-ai
15 Jul 2026
Model Releases

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

DGX agent

arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, t

model-releasesarxiv-cs-ai
10 Jul 2026
Safety

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

DGX agent

arXiv:2601.10504v2 Announce Type: replace Abstract: As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, rel

safetyarxiv-cs-cl
10 Jul 2026
Agents

From Legacy Documentation to OSCAL: An MCP-Based Agent Pipeline for Threat-Informed Continuous Compliance in Critical Infrastructure

DGX agent

arXiv:2607.08288v1 Announce Type: cross Abstract: In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed for risk as

agentsarxiv-cs-ai
10 Jul 2026
Agents

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

DGX agent

arXiv:2607.08282v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards creates sign

agentsarxiv-cs-ai
10 Jul 2026
Agents

MMAgent-R^2: Learning to Rerank and Reject for Agentic mRAG

DGX agent

arXiv:2607.07383v1 Announce Type: new Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic kn

agentsarxiv-cs-cv
9 Jul 2026
Agents

Non-contact, Real-time, Heart-rate Measurement using Image Processing with Commodity Cameras and AI Agents

DGX agent

arXiv:2607.06598v1 Announce Type: cross Abstract: Heart rate measurement is one of the key requirements for real-time health monitoring, in particular for health caring of elderly people. Traditional

agentsarxiv-cs-ai
9 Jul 2026
Model Releases

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

DGX agent

arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single 'pipe

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

DGX agent

arXiv:2607.07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

DGX agent

arXiv:2607.06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger rep

model-releasesarxiv-cs-ai
9 Jul 2026
Agents

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

DGX agent

arXiv:2606.14948v2 Announce Type: replace-cross Abstract: LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding is p

agentsarxiv-cs-ai
8 Jul 2026
Agents

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition

DGX agent

arXiv:2607.06256v1 Announce Type: new Abstract: Long-horizon household tasks require robots to compose many language-conditioned skills, yet the boundary between consecutive skills is rarely explicit.

agentsarxiv-cs-ro
8 Jul 2026
← Previous
1…7576777879…236
Next →