AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
Agents

Coachable agents for interactive gameplay

DGX agent

arXiv:2607.00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -

model-releasesarxiv-cs-cl
2 Jul 2026
Agents

Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization

DGX agent

arXiv:2603.26270v2 Announce Type: replace-cross Abstract: Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

DGX agent

arXiv:2607.01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. How

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

DGX agent

arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestr

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents

DGX agent

arXiv:2606.30970v1 Announce Type: new Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

DGX agent

arXiv:2606.31200v1 Announce Type: new Abstract: Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based me

model-releasesarxiv-cs-ai
1 Jul 2026
Agents

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

DGX agent

arXiv:2606.31134v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evad

agentsarxiv-cs-ai
1 Jul 2026
Model Releases

Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

DGX agent

arXiv:2606.31371v1 Announce Type: cross Abstract: When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

'For agents within an actual product, I have gravitated to using LangChain DeepAgents. I feel like their base harness is less heavy, and the…

DGX agent

'For agents within an actual product, I have gravitated to using LangChain DeepAgents. I feel like their base harness is less heavy, and their abstraction for building agents that use skills and tools

model-releasesharrison-chase--x
1 Jul 2026
Applications

Formal organizational structures are a useful way to think about the challenges of agents. They provide a template to thinking about how wor…

DGX agent

Formal organizational structures are a useful way to think about the challenges of agents. They provide a template to thinking about how work gets delegated up and down between smart expensive agents

applicationsethan-mollick--x
1 Jul 2026
Tutorials

Interrupt 2026 is coming to New York City and London this fall. → Hear from engineers + teams shipping agents in production → Keynotes from …

DGX agent

Interrupt 2026 is coming to New York City and London this fall. → Hear from engineers + teams shipping agents in production → Keynotes from @hwchase17 + industry leaders on what's next for agents → Ha

tutorialsharrison-chase--x
1 Jul 2026
Tools

“Valued work per watt is the score” @rolandgvc with the real details of self improving agents at @aiDotEngineer!

DGX agent

This post discusses a framework for evaluating self-improving AI agents where 'value per watt' serves as a key metric, likely examining how efficiently these agents generate useful output relative to

toolsswyx--x
1 Jul 2026
Safety

Verification-Gated Agentic Mission-State Governance for Intelligent Industrial Multi-Robot Systems

DGX agent

arXiv:2606.31339v1 Announce Type: new Abstract: Agentic artificial intelligence is increasingly used to decompose industrial tasks, propose robot actions, and adapt execution plans in dynamic cyber-ph

safetyarxiv-cs-ro
1 Jul 2026
Model Releases

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

DGX agent

arXiv:2606.29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observa

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation

DGX agent

arXiv:2606.28896v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) data augmentation is important for improving the generalization of data-driven SAR interpretation models, yet practical

agentsarxiv-cs-ai
30 Jun 2026
Safety

Agent Safety Is Action Alignment

DGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

safetyarxiv-cs-ai
30 Jun 2026
Agents

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

DGX agent

arXiv:2606.17555v2 Announce Type: replace-cross Abstract: Banks face two threat families with fundamentally different detection requirements: signature-based fraud (card-not-present attacks, account t

agentsarxiv-cs-ai
30 Jun 2026
Safety

Analytic Concept-Centric Memory for Agentic Embodied Manipulation

DGX agent

arXiv:2606.29774v1 Announce Type: new Abstract: Long-horizon embodied manipulation requires agents to remember persistent objects, track changing scene states, and reuse prior interaction knowledge. H

safetyarxiv-cs-ro
30 Jun 2026
Model Releases

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level cod…

DGX agent

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level code evolution. A Markdown harness becomes a project pack with

model-releasesdair-ai--x
30 Jun 2026
Agents

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

DGX agent

arXiv:2606.29746v1 Announce Type: new Abstract: Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains

agentsarxiv-cs-ai
30 Jun 2026
Safety

Entity Binding Failures in Tool-Augmented Agents

DGX agent

arXiv:2606.30531v1 Announce Type: new Abstract: Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requeste

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

Have your agent record video demos of its work with shot-scraper video

DGX agent

shot-scraper video is a new command introduced in today's shot-scraper 1.10 release which accepts a storyboard.yml file defining a routine to run against a web application and uses Playwright to recor

model-releasessimon-willison
30 Jun 2026
Model Releases

LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard

DGX agent

arXiv:2606.30005v1 Announce Type: new Abstract: Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management age

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vecto…

DGX agent

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vectors distilled from contrastive historical experience. Textual

model-releasesdair-ai--x
30 Jun 2026
Agents

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

DGX agent

arXiv:2606.29932v1 Announce Type: new Abstract: Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and

agentsarxiv-cs-ai
30 Jun 2026
Tutorials

SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents

DGX agent

arXiv:2606.28434v1 Announce Type: cross Abstract: Long-horizon software engineering agents often need to manage lengthy and noisy interaction histories under limited context budgets. Existing memory m

tutorialsarxiv-cs-ai
30 Jun 2026
Model Releases

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing o…

DGX agent

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing our journey and walking you through building your first agent

model-releasesswyx--x
30 Jun 2026
Agents

TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging

DGX agent

arXiv:2606.29763v1 Announce Type: cross Abstract: Topological data analysis (TDA), particularly persistent homology (PH), captures geometric structural properties in medical images (e.g., connected co

agentsarxiv-cs-ai
30 Jun 2026
Agents

Agent confidence on the technical frontier

DGX agent

Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mount

agentsmit-tech-review
29 Jun 2026
Agents

Agentic Episodic Control

DGX agent

arXiv:2506.01442v2 Announce Type: replace Abstract: Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attempt to alle

agentsarxiv-cs-ai
29 Jun 2026
Model Releases

Agentic Hardware Design as Repository-Level Code Evolution

DGX agent

arXiv:2606.28279v1 Announce Type: cross Abstract: We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled int

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

From Detection to Action: Using LLM Agents for Fault-Tolerant Control

DGX agent

arXiv:2606.28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constr

model-releasesarxiv-cs-lg
29 Jun 2026
Safety

Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs

DGX agent

arXiv:2601.21233v2 Announce Type: replace Abstract: Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-d

safetyarxiv-cs-ai
29 Jun 2026
Hardware

One day left to submit your projects for the Hermes Agent Accelerated Business Hackathon presented by @NVIDIAAI × @stripe × @NousResearch! S…

DGX agent

One day left to submit your projects for the Hermes Agent Accelerated Business Hackathon presented by @NVIDIAAI × @stripe × @NousResearch! Submissions close 11:59 PM PT tomorrow, June 30th. Last minut

hardwarenous-research--x
29 Jun 2026
Safety

Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems

DGX agent

arXiv:2505.23847v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint disas

safetyarxiv-cs-ai
29 Jun 2026
Model Releases

Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents

DGX agent

arXiv:2606.27472v1 Announce Type: cross Abstract: Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revise

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

DGX agent

arXiv:2606.27669v1 Announce Type: new Abstract: Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval a

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

DGX agent

arXiv:2606.26216v1 Announce Type: cross Abstract: We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability det

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

DGX agent

arXiv:2606.21649v2 Announce Type: replace Abstract: Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This

agentsarxiv-cs-cl
26 Jun 2026
Model Releases

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

DGX agent

arXiv:2606.26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-doma

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

DGX agent

arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tas

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

Prompt, Plan, Extract: Zero-Shot Agentic LLMs Workflows for Lung Pathology Extraction from Clinical Narratives

DGX agent

arXiv:2606.19852v2 Announce Type: replace Abstract: Information extraction from pathology reports is essential for cancer staging, tumor registry population. Yet key data remains embedded in narrative

agentsarxiv-cs-cl
26 Jun 2026
Agents

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

DGX agent

arXiv:2606.26722v1 Announce Type: new Abstract: The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypot

agentsarxiv-cs-ai
26 Jun 2026
Safety

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

DGX agent

arXiv:2606.25421v1 Announce Type: new Abstract: Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. Howeve

safetyarxiv-cs-cl
25 Jun 2026
Safety

GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning

DGX agent

arXiv:2606.25073v1 Announce Type: new Abstract: In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch fo

safetyarxiv-cs-lg
25 Jun 2026
Agents

MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models

DGX agent

arXiv:2606.24815v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging robotic control systems that integrate perception, language understanding, and action generation in a

agentsarxiv-cs-ro
25 Jun 2026
Agents

Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

DGX agent

arXiv:2606.25361v1 Announce Type: new Abstract: Prior research on memory mechanism in RAG-based conversational system has emphasized how memory is stored and retrieved. However, far less is known abou

agentsarxiv-cs-cl
25 Jun 2026
← Previous
1…116117118119120…375
Next →