AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
10 Jul 2026

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments

Model ReleasesDGX agent

arXiv:2603.16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in d

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents

Model ReleasesDGX agent

arXiv:2607.08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and

9 Jul 2026

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2607.06854v1 Announce Type: cross Abstract: Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade,

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

Model ReleasesDGX agent

arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the

End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

AgentsDGX agent

arXiv:2607.06964v1 Announce Type: cross Abstract: Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) a

Reliable and Developer-Aligned Evaluation of Agents for Software Engineering

AgentsDGX agent

arXiv:2607.06713v1 Announce Type: cross Abstract: Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contribut

8 Jul 2026

Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents

Model ReleasesDGX agent

arXiv:2510.19771v4 Announce Type: replace Abstract: LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and so

CHARLIE: An On-Premise Multi-Agent Retrieval-Augmented Generation System for Evidential Reasoning in Forensic Science

Local AiDGX agent

arXiv:2607.05428v1 Announce Type: cross Abstract: We present Charlie, an on-premise multi-agent Retrieval-Augmented Generation (RAG) system for structured evidential processing in digital forensic env

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Model ReleasesDGX agent

arXiv:2607.06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantia

Ex-GitHub chief’s Entire opens distributed Git network for the AI agent era

AgentsDGX agent

Entire Inc., the developer-platform startup founded by former GitHub Chief Executive Thomas Dohmke, today launched a preview of a distributed Git network built to let artificial intelligence coding ag

From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b

AgentsDGX agent

arXiv:2607.06452v1 Announce Type: cross Abstract: Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidenc

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

HardwareDGX agent

arXiv:2607.05511v1 Announce Type: new Abstract: Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. Howe

// Memory becomes an action space // Great paper on long-term memory for agents. (bookmark it) In short, it's discusses the use of a learned…

SafetyDGX agent

// Memory becomes an action space // Great paper on long-term memory for agents. (bookmark it) In short, it's discusses the use of a learned policy for using memory at the right granularity. Most memo

Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms

AgentsDGX agent

arXiv:2607.06489v1 Announce Type: new Abstract: The dairy industry in Ireland has a large potential for the integration of renewable energy and the reduction of carbon emissions. However, researchers

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

Model ReleasesDGX agent

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tuned its Deep Agents h

PolyJarvis: An LLM-Orchestrated Agent for Automated All-Atom Molecular Dynamics of Amorphous Homopolymers

AgentsDGX agent

arXiv:2604.02537v2 Announce Type: replace Abstract: All-atom molecular dynamics (MD) simulations can predict polymer properties from molecular structure, yet their execution requires specialized exper

Prompt-to-Paper: Agentic AI System for Bioinformatics

AgentsDGX agent

arXiv:2607.05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical defi

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

Model ReleasesDGX agent

arXiv:2607.06411v1 Announce Type: cross Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of

SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

Model ReleasesDGX agent

arXiv:2606.13757v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is mer

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their nav

7 Jul 2026

Agentic Very Long Video Understanding

AgentsDGX agent

arXiv:2601.18157v3 Announce Type: replace Abstract: The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual underst

AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression

AgentsDGX agent

arXiv:2512.13956v4 Announce Type: replace-cross Abstract: Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and met

AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation

Model ReleasesDGX agent

arXiv:2607.02520v1 Announce Type: cross Abstract: Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether gener

EEG-SpikeAgent: Agentic Closed-Loop Program Synthesis for Automated EEG Spike Detection

AgentsDGX agent

arXiv:2607.04558v1 Announce Type: cross Abstract: Automated detection of interictal epileptiform discharges in scalp electroencephalography (EEG) is clinically important, but recent high-performing de

Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning

AgentsDGX agent

arXiv:2607.02588v1 Announce Type: cross Abstract: Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

Model ReleasesDGX agent

arXiv:2606.20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under

PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents

Model ReleasesDGX agent

arXiv:2607.04089v1 Announce Type: new Abstract: Lifelong agents need more than larger context windows and better retrieval. They need memories that can persist, evolve, and be corrected without forcin

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe

Model ReleasesDGX agent

arXiv:2607.03451v1 Announce Type: cross Abstract: While skill optimization for autonomous agents has gained traction, existing methods rely on complex pipelines. This leaves a fundamental question una

6 Jul 2026

Shift into high gear with agents: Securing the software-defined vehicle

Model ReleasesDGX agent

The automotive industry is at a pivotal crossroads as it hits the gas on adopting new technology. The era of the traditional connected vehicle has shifted into the age of the software-defined vehicle

When it comes to Chess, @viditchess is an absolute legend. Recently, he entered the Hermes Agent Accelerated Business Hackathon to see who c…

Model ReleasesDGX agent

When it comes to Chess, @viditchess is an absolute legend. Recently, he entered the Hermes Agent Accelerated Business Hackathon to see who could build the best agent using Nemotron. Here's what he sub

4 Jul 2026

We've created a comprehensive Retrieval Harness for modern agentic retrieval in 2026. The harness provides a persistent data pipeline that c…

Model ReleasesDGX agent

We've created a comprehensive Retrieval Harness for modern agentic retrieval in 2026. The harness provides a persistent data pipeline that can connect to a data source, index and update a large knowle

3 Jul 2026

CausalSteward: An Agentic Divide-Conquer-Combine Copilot for Causal Discovery

AgentsDGX agent

arXiv:2607.01936v1 Announce Type: cross Abstract: Learning causal models from high-dimensional data is a significant challenge, particularly in real-world settings where violations of core assumptions

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Model ReleasesDGX agent

arXiv:2606.14249v2 Announce Type: replace Abstract: AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model obs

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

Model ReleasesDGX agent

arXiv:2604.04532v2 Announce Type: replace-cross Abstract: Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language

2 Jul 2026

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Model ReleasesDGX agent

arXiv:2607.01084v1 Announce Type: new Abstract: While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynami

Coachable agents for interactive gameplay

AgentsDGX agent

arXiv:2607.00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

Model ReleasesDGX agent

arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -

Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization

AgentsDGX agent

arXiv:2603.26270v2 Announce Type: replace-cross Abstract: Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Model ReleasesDGX agent

arXiv:2607.01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. How

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Model ReleasesDGX agent

arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestr

1 Jul 2026

AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2606.30970v1 Announce Type: new Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

Model ReleasesDGX agent

arXiv:2606.31200v1 Announce Type: new Abstract: Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based me

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

AgentsDGX agent

arXiv:2606.31134v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evad

Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

Model ReleasesDGX agent

arXiv:2606.31371v1 Announce Type: cross Abstract: When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned

'For agents within an actual product, I have gravitated to using LangChain DeepAgents. I feel like their base harness is less heavy, and the…

Model ReleasesDGX agent

'For agents within an actual product, I have gravitated to using LangChain DeepAgents. I feel like their base harness is less heavy, and their abstraction for building agents that use skills and tools

Formal organizational structures are a useful way to think about the challenges of agents. They provide a template to thinking about how wor…

ApplicationsDGX agent

Formal organizational structures are a useful way to think about the challenges of agents. They provide a template to thinking about how work gets delegated up and down between smart expensive agents

Interrupt 2026 is coming to New York City and London this fall. → Hear from engineers + teams shipping agents in production → Keynotes from …

TutorialsDGX agent

Interrupt 2026 is coming to New York City and London this fall. → Hear from engineers + teams shipping agents in production → Keynotes from @hwchase17 + industry leaders on what's next for agents → Ha

“Valued work per watt is the score” @rolandgvc with the real details of self improving agents at @aiDotEngineer!

ToolsDGX agent

This post discusses a framework for evaluating self-improving AI agents where 'value per watt' serves as a key metric, likely examining how efficiently these agents generate useful output relative to

Verification-Gated Agentic Mission-State Governance for Intelligent Industrial Multi-Robot Systems

SafetyDGX agent

arXiv:2606.31339v1 Announce Type: new Abstract: Agentic artificial intelligence is increasingly used to decompose industrial tasks, propose robot actions, and adapt execution plans in dynamic cyber-ph

30 Jun 2026

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

Model ReleasesDGX agent

arXiv:2606.29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observa

A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation

AgentsDGX agent

arXiv:2606.28896v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) data augmentation is important for improving the generalization of data-driven SAR interpretation models, yet practical

Agent Safety Is Action Alignment

SafetyDGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

AgentsDGX agent

arXiv:2606.17555v2 Announce Type: replace-cross Abstract: Banks face two threat families with fundamentally different detection requirements: signature-based fraud (card-not-present attacks, account t

Analytic Concept-Centric Memory for Agentic Embodied Manipulation

SafetyDGX agent

arXiv:2606.29774v1 Announce Type: new Abstract: Long-horizon embodied manipulation requires agents to remember persistent objects, track changing scene states, and reuse prior interaction knowledge. H

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level cod…

Model ReleasesDGX agent

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level code evolution. A Markdown harness becomes a project pack with

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

AgentsDGX agent

arXiv:2606.29746v1 Announce Type: new Abstract: Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains

Entity Binding Failures in Tool-Augmented Agents

SafetyDGX agent

arXiv:2606.30531v1 Announce Type: new Abstract: Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requeste

Have your agent record video demos of its work with shot-scraper video

Model ReleasesDGX agent

shot-scraper video is a new command introduced in today's shot-scraper 1.10 release which accepts a storyboard.yml file defining a routine to run against a web application and uses Playwright to recor

LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard

Model ReleasesDGX agent

arXiv:2606.30005v1 Announce Type: new Abstract: Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management age

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vecto…

Model ReleasesDGX agent

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vectors distilled from contrastive historical experience. Textual

← Previous
1…9293949596…300
Next →