AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
Safety

Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents

DGX agent

arXiv:2505.11708v3 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) agents are increasingly used to simulate sophisticated cyberattacks, but their decision-making processes remain op

safetyarxiv-cs-lg
18 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Verifiable Agentic Infrastructure: Proof-Derived Authorization for Sovereign AI Systems

DGX agent

arXiv:2605.15228v1 Announce Type: new Abstract: Modern cloud and enterprise systems rely on identity-centric authorization, assuming that callers possessing valid credentials are safe to execute comma

safetyarxiv-cs-ai
18 May 2026
Agents

Articraft: An Agentic System for Scalable Articulated 3D Asset Generation

DGX agent

arXiv:2605.15187v1 Announce Type: new Abstract: A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large lan

agentsarxiv-cs-cv
15 May 2026
Agents

CA2: Code-Aware Agent for Automated Game Testing

DGX agent

arXiv:2605.13918v1 Announce Type: cross Abstract: Automated game testing is important for verifying game functionality, but it remains a costly and time-consuming process. Manual testing often misses

agentsarxiv-cs-lg
15 May 2026
Agents

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making

DGX agent

arXiv:2605.14403v1 Announce Type: new Abstract: Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (

agentsarxiv-cs-cv
15 May 2026
Model Releases

FutureSim: Replaying World Events to Evaluate Adaptive Agents

DGX agent

arXiv:2605.15188v1 Announce Type: cross Abstract: AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently m

model-releasesarxiv-cs-ai
15 May 2026
Local Ai

It has been a pleasure collaborating with the @NVIDIAAI team to ensure that Hermes Agent runs perfectly on DGX Spark!

DGX agent

It has been a pleasure collaborating with the @NVIDIAAI team to ensure that Hermes Agent runs perfectly on DGX Spark! Run @NousResearch's Hermes Agent fully locally on DGX Spark. 🚀 Our newest playbook

local-ainous-research--x
15 May 2026
Model Releases

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

DGX agent

arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r

model-releasesarxiv-cs-cl
15 May 2026
Agents

Nexus : An Agentic Framework for Time Series Forecasting

DGX agent

arXiv:2605.14389v1 Announce Type: new Abstract: Time series forecasting is not just numerical extrapolation, but often requires reasoning with unstructured contextual data such as news or events. Whil

agentsarxiv-cs-ai
15 May 2026
Model Releases

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

DGX agent

arXiv:2605.14002v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into op

model-releasesarxiv-cs-ai
15 May 2026
Agents

SR-Platform: An Agentic Pipeline for Natural Language-Driven Robot Simulation Environment Synthesis

DGX agent

arXiv:2605.14700v1 Announce Type: new Abstract: Generating robot simulation environments remains a major bottleneck in simulation-based robot learning. Constructing a training-ready MuJoCo scene typic

agentsarxiv-cs-ro
15 May 2026
Safety

TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

DGX agent

arXiv:2605.13909v1 Announce Type: cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonic

safetyarxiv-cs-ai
15 May 2026
Research

Continual Harness: Online Adaptation for Self-Improving Foundation Agents [R]

DGX agent

Continual Harness proposes an approach to online adaptation for foundation agents that moves beyond traditional gradient-based retraining by introducing a dual-agent architecture (Teacher/Student) wit

researchr-machinelearning
14 May 2026
Model Releases

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents

DGX agent

arXiv:2601.18842v3 Announce Type: replace-cross Abstract: As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive informati

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Kimi K2.6 is now open-weight #1 on Finance Agent Benchmark V2.

DGX agent

Kimi K2.6 is now open-weight #1 on Finance Agent Benchmark V2. Can AI do the job of a financial analyst? We just released V2 of our Finance Agent Benchmark and tested the frontier models. The results

model-releaseskimi-moonshot--x
14 May 2026
Model Releases

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

DGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

model-releasesarxiv-cs-cv
14 May 2026
Agents

Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue

DGX agent

arXiv:2605.12856v1 Announce Type: new Abstract: The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with {em malicious intent} may

agentsarxiv-cs-ai
14 May 2026
Agents

Multi-Agent Systems in Emergency Departments: Validation Study on a ED Digital Twin

DGX agent

arXiv:2605.13345v1 Announce Type: new Abstract: Emergency departments (ED) face challenges in patient care and resource management. We propose to explore optimization strategies in a realistic and fle

agentsarxiv-cs-ai
14 May 2026
Agents

RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records

DGX agent

arXiv:2507.15867v2 Announce Type: replace-cross Abstract: Rare diseases affect 1 in 10 Americans yet remain systematically underdocumented in clinical records. ICD-based systems cannot capture their b

agentsarxiv-cs-ai
14 May 2026
Model Releases

RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

DGX agent

arXiv:2605.13391v1 Announce Type: new Abstract: The rise of multi-modal large language models (MLLMs) is shifting remote sensing (RS) intelligence from 'see' to 'action', as OpenClaw-style frameworks

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

The Building AI Agents with MongoDB and LangGraph Skill Badge can be yours on May 28th. Here's how you earn it 👉 Join our LIVE workshop, bu…

DGX agent

The Building AI Agents with MongoDB and LangGraph Skill Badge can be yours on May 28th. Here's how you earn it 👉 Join our LIVE workshop, build a real agent with MongoDB, Claude Sonnet, and LangGraph,

model-releasesharrison-chase--x
14 May 2026
Agents

What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

DGX agent

arXiv:2605.12745v1 Announce Type: cross Abstract: Discrepancies between an agent's actual knowledge and what a person thinks the agent knows can hinder interactions. If an agent could detect such disc

agentsarxiv-cs-ai
14 May 2026
Model Releases

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

DGX agent

arXiv:2508.07642v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3

model-releasesarxiv-cs-cl
13 May 2026
Agents

From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation

DGX agent

arXiv:2605.11951v1 Announce Type: new Abstract: Although robotic manipulation has made significant progress, reliable execution remains challenging because task failures are inevitable in dynamic and

agentsarxiv-cs-ro
13 May 2026
Local Ai

Hermes Agent now runs natively on NVIDIA RTX PCs and DGX Spark. Hermes is designed for exactly the kind of always-on workload that NVIDIA's …

DGX agent

Hermes Agent now runs natively on NVIDIA RTX PCs and DGX Spark. Hermes is designed for exactly the kind of always-on workload that NVIDIA's hardware is built for, and their blog explains in depth why

local-ainous-research--x
13 May 2026
Safety

Learning Agentic Policy from Action Guidance

DGX agent

arXiv:2605.12004v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training si

safetyarxiv-cs-cl
13 May 2026
Agents

Microsoft’s new agentic security system MDASH uncovers four critical Windows RCE flaws

DGX agent

Microsoft Corp. today detailed a new artificial intelligence-powered vulnerability discovery system that uncovered 16 previously unknown flaws in Windows networking and authentication components, incl

agentssiliconangle
13 May 2026
Safety

SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs

DGX agent

arXiv:2605.12039v1 Announce Type: new Abstract: Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entr

safetyarxiv-cs-cl
13 May 2026
Agents

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

DGX agent

arXiv:2605.08956v1 Announce Type: new Abstract: A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they alread

agentsarxiv-cs-ai
12 May 2026
Model Releases

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

DGX agent

arXiv:2605.10787v1 Announce Type: new Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the 'last mile' of commercial software automation. In real-world scenarios,

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

DGX agent

arXiv:2605.09998v1 Announce Type: cross Abstract: Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CrackMeBench: Binary Reverse Engineering for Agents

DGX agent

arXiv:2605.10597v1 Announce Type: cross Abstract: Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-fl

model-releasesarxiv-cs-ai
12 May 2026
Safety

Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents

DGX agent

arXiv:2605.08717v1 Announce Type: cross Abstract: Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad

safetyarxiv-cs-ai
12 May 2026
Model Releases

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

DGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents

DGX agent

arXiv:2605.10332v1 Announce Type: new Abstract: Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied enviro

model-releasesarxiv-cs-ai
12 May 2026
Agents

Generating synthetic electronic health record data using agent-based models to evaluate machine learning robustness under mass casualty incidents

DGX agent

arXiv:2605.09951v1 Announce Type: new Abstract: ML models in healthcare are typically evaluated using curated real-world EHR data. A key limitation of such evaluations is that they may fail to assess

agentsarxiv-cs-lg
12 May 2026
Agents

HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics

DGX agent

arXiv:2507.15518v5 Announce Type: replace Abstract: Creating an immersive and interactive theatrical experience is a long-term goal in the field of interactive narrative. The emergence of large langua

agentsarxiv-cs-ai
12 May 2026
Research

How Mobile World Model Guides GUI Agents?

DGX agent

arXiv:2605.10347v1 Announce Type: new Abstract: Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable predi

researcharxiv-cs-ai
12 May 2026
Safety

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization

DGX agent

arXiv:2605.08978v1 Announce Type: new Abstract: Recent advancements in agentic test-time scaling allow models to gather environmental feedback before committing to final actions. A key limitation of e

safetyarxiv-cs-ai
12 May 2026
Model Releases

Log analysis is necessary for credible evaluation of AI agents

DGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

model-releasesarxiv-cs-ai
12 May 2026
Agents

Manifold scores 7,700 MCP servers in Manifest expansion aimed at agent security teams

DGX agent

Artificial intelligence detection and response platform startup Manifold Security Inc. today announced an expansion of its Manifest supply chain intelligence tool to cover Model Context Protocol serve

agentssiliconangle
12 May 2026
Model Releases

MDGYM: Benchmarking AI Agents on Molecular Simulations

DGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

DGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

model-releasesarxiv-cs-ai
12 May 2026
Hardware

Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts

DGX agent

arXiv:2605.09055v1 Announce Type: cross Abstract: Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-

hardwarearxiv-cs-ai
12 May 2026
Model Releases

Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning

DGX agent

arXiv:2605.09822v1 Announce Type: cross Abstract: We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use p

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

DGX agent

arXiv:2605.08876v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that execute tool-augmented, multi-step tasks, where latency is a critical f

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

DGX agent

arXiv:2604.01527v3 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fi

model-releasesarxiv-cs-ai
12 May 2026
Safety

Route by State, Recover from Trace: STAR with Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning

DGX agent

arXiv:2605.10057v1 Announce Type: new Abstract: Compositional spatiotemporal reasoning often requires a system to invoke multiple heterogeneous specialists, such as geometric, temporal, topological, a

safetyarxiv-cs-ai
12 May 2026
← Previous
1…121122123124125…375
Next →