AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
24 Jul 2026

Opus 5 now available in Hermes Agent

Model ReleasesDGX agent

Claude Opus 5 is now released in the Hermes Agent, a product of Nous Research and Teknium. Users can access the model through multiple gateways, including the Nous Portal, OpenRouter, and Anthropic Di

Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

SafetyDGX agent

arXiv:2607.20708v1 Announce Type: new Abstract: A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that Phi_r grow

Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

AgentsDGX agent

arXiv:2607.20431v1 Announce Type: new Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet conv

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

Model ReleasesDGX agent

arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Te

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Model ReleasesDGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

23 Jul 2026

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Model ReleasesDGX agent

arXiv:2607.14573v3 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows,

Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving

AgentsDGX agent

arXiv:2607.19774v1 Announce Type: new Abstract: Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collab

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

Model ReleasesDGX agent

arXiv:2607.19038v1 Announce Type: new Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

Model ReleasesDGX agent

arXiv:2607.14186v4 Announce Type: replace-cross Abstract: Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

SafetyDGX agent

arXiv:2607.19351v1 Announce Type: new Abstract: LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions throu

Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation

AgentsDGX agent

arXiv:2607.19767v1 Announce Type: new Abstract: A rich and recognizable component library is the cornerstone of printed circuit board (PCB) design and generation. Traditionally, engineers manually cre

The Ethics of Autonomous AI Agents for Offensive Security

SafetyDGX agent

arXiv:2607.20255v1 Announce Type: cross Abstract: LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and o

22 Jul 2026

browser-search v2.0 — From the balaclava to the badge: your agent now browses everywhere

Model ReleasesDGX agent

Today an AI agent trying to browse the web is like a thief in a balaclava sneaking around a police academy. Site protections block it, challenge it, turn it away. browser-search flips the script: your

21 Jul 2026

A Fireside Chat with Cat and Thariq from the Claude Code team

Model ReleasesDGX agent

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, c

Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

HardwareDGX agent

NVIDIA’s Rubin GPU, the core of the Vera Rubin platform, delivers up to 10× the agentic inference throughput per watt compared with previous generations, using 336 billion transistors, 224 SMs, 896 Te

Now in preview: Find and fix software vulnerabilities with CodeMender

Model ReleasesDGX agent

As adversarial AI threats accelerate attacks on code, security teams must counter them with machine-speed defenses that can automate code remediation and fight AI with AI. CodeMender is our managed co

16 Jul 2026

EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

Model ReleasesDGX agent

arXiv:2607.13792v1 Announce Type: new Abstract: Most daily activities are inherently procedural. However, existing evaluations for egocentric video understanding seldom address procedural understandin

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

AgentsDGX agent

arXiv:2607.13679v1 Announce Type: new Abstract: AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen o

15 Jul 2026

Audio perception layer for LLM agents, with a memory that grows through use

Model ReleasesDGX agent

LLMs handle speech well once you run speech-to-text. They don't hear the rest: a bird outside, a glass breaking two rooms away, a smoke alarm two floors down. I've been working on an experimental open

Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition

Local AiDGX agent

arXiv:2607.12911v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used for dietary assessment from meal images, where retrieval-augmented grounding was shown to

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

AgentsDGX agent

arXiv:2607.11019v2 Announce Type: replace Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineerin

Unveiling Complex Collective Behaviors from Simple Rewards

AgentsDGX agent

arXiv:2607.12861v1 Announce Type: cross Abstract: Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic an

14 Jul 2026

Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills

HardwareDGX agent

NVIDIA Cosmos 3 was post‑trained in under a day using TAO agent skills and LoRA adapters, raising accuracy on the Woven Traffic Safety video QA dataset from 54.41 % to 93.35 %. The mixture‑of‑transfor

13 Jul 2026

Livestream Alert: Run ComfyUI From Claude/Cursor with Comfy MCP Host: @PurzBeats Comfy MCP lets Claude, Cursor, Amp and almost any AI agent …

Model ReleasesDGX agent

Livestream Alert: Run ComfyUI From Claude/Cursor with Comfy MCP Host: @PurzBeats Comfy MCP lets Claude, Cursor, Amp and almost any AI agent you're already using build, run, and iterate real Comfy Clou

10 Jul 2026

Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies

AgentsDGX agent

This paper was accepted at the AI4TCI (Workshop on AI for Secure and Trustworthy Critical Infrastructure Systems) Workshop at the International Conference on Availability, Reliability and Security (AR

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

Model ReleasesDGX agent

arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, t

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

SafetyDGX agent

arXiv:2601.10504v2 Announce Type: replace Abstract: As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, rel

From Legacy Documentation to OSCAL: An MCP-Based Agent Pipeline for Threat-Informed Continuous Compliance in Critical Infrastructure

AgentsDGX agent

arXiv:2607.08288v1 Announce Type: cross Abstract: In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed for risk as

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

AgentsDGX agent

arXiv:2607.08282v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards creates sign

9 Jul 2026

MMAgent-R^2: Learning to Rerank and Reject for Agentic mRAG

AgentsDGX agent

arXiv:2607.07383v1 Announce Type: new Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic kn

Non-contact, Real-time, Heart-rate Measurement using Image Processing with Commodity Cameras and AI Agents

AgentsDGX agent

arXiv:2607.06598v1 Announce Type: cross Abstract: Heart rate measurement is one of the key requirements for real-time health monitoring, in particular for health caring of elderly people. Traditional

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Model ReleasesDGX agent

arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single 'pipe

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2607.07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Model ReleasesDGX agent

arXiv:2607.06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger rep

8 Jul 2026

Automating cross-repo documentation with GitHub Agentic Workflows

AgentsDGX agent

Explore how the Aspire team turns merged product changes into SME-reviewed docs pull requests, closing the gap between release and documentation. The post Automating cross-repo documentation with GitH

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

AgentsDGX agent

arXiv:2606.14948v2 Announce Type: replace-cross Abstract: LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding is p

Data for Agents

ToolsDGX agent

This resource discusses open datasets and data resources designed to support the development and training of AI agents, likely covering data collection methodologies, benchmark datasets, and best prac

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition

AgentsDGX agent

arXiv:2607.06256v1 Announce Type: new Abstract: Long-horizon household tasks require robots to compose many language-conditioned skills, yet the boundary between consecutive skills is rarely explicit.

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports

AgentsDGX agent

arXiv:2607.06435v1 Announce Type: new Abstract: Data from Singapore indicated that about 31% of the population had evidence of Helicobacter pylori infection. Persistent H. pylori infection is associat

How Personas Can Influence Agents to Play Split or Steal

TutorialsDGX agent

arXiv:2607.05398v1 Announce Type: new Abstract: Personas are often employed to guide large language model agents, yet their effectiveness in shaping strategic behavior in social dilemma settings remai

Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test

Model ReleasesDGX agent

arXiv:2607.06001v1 Announce Type: new Abstract: We report a pre-registered, two-part experiment on small economies of frontier language-model agents (Claude Opus 4.8), testing two quantitative predict

LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability

Model ReleasesDGX agent

arXiv:2607.06157v1 Announce Type: cross Abstract: Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach a

Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agents

Model ReleasesDGX agent

arXiv:2607.05690v1 Announce Type: new Abstract: Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn. We study the

Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?

SafetyDGX agent

arXiv:2508.01109v3 Announce Type: replace Abstract: We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery (capturin

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

AgentsDGX agent

arXiv:2607.06074v1 Announce Type: cross Abstract: Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped

StepShield: When, Not Whether to Intervene on Rogue Agents

Model ReleasesDGX agent

arXiv:2601.22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We in

TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting

AgentsDGX agent

arXiv:2607.06349v1 Announce Type: new Abstract: Building sensors are embedded in physical topology, spatial hierarchy, and operational context, yet existing forecasters often treat them as isolated ti

7 Jul 2026

ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning

SafetyDGX agent

arXiv:2607.04162v1 Announce Type: cross Abstract: Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

Local AiDGX agent

arXiv:2607.03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effective

AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection

AgentsDGX agent

arXiv:2603.23115v2 Announce Type: replace Abstract: The realism of AI-generated images (AIGI) poses increasing challenges for reliable forensic detection, where heterogeneous expert detectors may prod

AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

Model ReleasesDGX agent

arXiv:2607.04758v1 Announce Type: new Abstract: Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation re

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI

SafetyDGX agent

arXiv:2607.03510v1 Announce Type: cross Abstract: Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and retrieval-aug

CoGen3D: An Agentic Human-AI Co-Design Pipeline for 3D Asset Generation for Virtual Reality

AgentsDGX agent

arXiv:2607.03731v1 Announce Type: cross Abstract: Creating 3D assets for virtual reality requires modeling expertise, which restricts the authorship of immersive experiences. Existing generative AI to

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

AgentsDGX agent

arXiv:2607.05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to th

CyberProof launches agentic MXDR to automate two-thirds of investigations

Model ReleasesDGX agent

CyberProof Inc., a co-managed security services company owned by UST Global Inc., today launched an agentic artificial intelligence service that it says can autonomously handle up to two-thirds of sec

Object-Centric Environment Modeling for Agentic Tasks

Local AiDGX agent

arXiv:2607.02846v1 Announce Type: new Abstract: Large language model (LLM) agents can improve through accumulated experience, but free-form textual memories become difficult to maintain, validate, and

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.04242v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

AgentsDGX agent

arXiv:2607.05382v1 Announce Type: cross Abstract: Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tai

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

Model ReleasesDGX agent

arXiv:2510.19186v3 Announce Type: replace Abstract: Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and to

6 Jul 2026

Contextual Policies in Omnigent: Using session state to better govern AI agents

IndustryDGX agent

Databricks describes contextual policies in Omnigent, a framework that leverages session state to implement more effective governance and control mechanisms for AI agents. By maintaining and utilizing

← Previous
1…102103104105106…300
Next →