AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Applications

Context as Prior: Bayesian-Inspired Intent Inference for Non-Speaking Agents with a Household Cat Testbed

DGX agent

arXiv:2604.27445v1 Announce Type: new Abstract: Many agents in real-world environments cannot reliably communicate their goals through language, including household pets, pre-verbal infants, and other

applicationsarxiv-cs-cv
1 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

DeepTutor: Towards Agentic Personalized Tutoring

DGX agent

arXiv:2604.26962v1 Announce Type: cross Abstract: Education represents one of the most promising real-world applications for Large Language Models (LLMs). However, conventional tutoring systems rely o

model-releasesarxiv-cs-ai
1 May 2026
Safety

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

DGX agent

arXiv:2604.27045v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memor

safetyarxiv-cs-ai
1 May 2026
Agents

End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians

DGX agent

arXiv:2604.27309v1 Announce Type: new Abstract: Clinical AI systems require not just point-in-time evaluation but continuous governance: the ongoing practice of monitoring, evaluating, iterating, and

agentsarxiv-cs-ai
1 May 2026
Agents

Heterogeneous Scientific Foundation Model Collaboration

DGX agent

arXiv:2604.27351v1 Announce Type: new Abstract: Agentic large language model systems have demonstrated strong capabilities. However, their reliance on language as the universal interface fundamentally

agentsarxiv-cs-ai
1 May 2026
Safety

Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents

DGX agent

arXiv:2604.27283v1 Announce Type: cross Abstract: Large language model (LLM)-based coding agents increasingly rely on external memory to reuse prior debugging experience, repair traces, and repository

safetyarxiv-cs-ai
1 May 2026
Agents

Optimal Stop-Loss and Take-Profit Parameterization for Autonomous Trading Agent Swarm

DGX agent

arXiv:2604.27150v1 Announce Type: new Abstract: Autonomous crypto trading systems often spend most of their design effort on finding entries, while exits are left to fixed rules that are rarely tested

agentsarxiv-cs-ai
1 May 2026
Model Releases

Progressive Multi-Agent Reasoning for Biological Perturbation Prediction

DGX agent

arXiv:2602.07408v2 Announce Type: replace Abstract: Predicting gene regulation responses to biological perturbations requires reasoning about underlying biological causalities. While large language mo

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

DGX agent

arXiv:2604.28093v1 Announce Type: new Abstract: Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the

model-releasesarxiv-cs-ai
1 May 2026
Safety

Evaluating Strategic Reasoning in Forecasting Agents

DGX agent

arXiv:2604.26106v1 Announce Type: new Abstract: Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to t

safetyarxiv-cs-ai
30 Apr 2026
Model Releases

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents

DGX agent

arXiv:2511.02399v2 Announce Type: replace-cross Abstract: Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requireme

model-releasesarxiv-cs-ai
30 Apr 2026
Agents

Lightweight Quantum Agent for Edge Systems: Joint PQC and NOMA Resource Allocation

DGX agent

arXiv:2604.25980v1 Announce Type: cross Abstract: In the context of quantum secure scenarios, existing research on mobile edge devices and intelligent computing and edge (ICE) systems based on the Non

agentsarxiv-cs-ai
30 Apr 2026
Local Ai

Provable Coordination for LLM Agents via Message Sequence Charts

DGX agent

arXiv:2604.17612v2 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are difficult to reason about. Coordination errors such as deadlocks or type-mismatc

local-aiarxiv-cs-ai
30 Apr 2026
Agents

StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

DGX agent

arXiv:2508.01875v4 Announce Type: replace Abstract: Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offl

agentsarxiv-cs-cv
30 Apr 2026
Model Releases

Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver

DGX agent

arXiv:2604.25067v1 Announce Type: cross Abstract: Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks mea

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

DGX agent

arXiv:2604.24929v1 Announce Type: new Abstract: Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editin

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Evaluating the Search Agent in a Parallel World

DGX agent

arXiv:2603.04751v2 Announce Type: replace Abstract: Integrating web search tools has significantly extended the capability of LLMs to address open-world, real-time, and long-tail problems. However, ev

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

GradMAP: Gradient-Based Multi-Agent Proximal Learning for Grid-Edge Flexibility

DGX agent

arXiv:2604.24549v1 Announce Type: cross Abstract: Coordinating large populations of grid-edge devices requires learning methods that remain fully decentralised in deployment while still respecting thr

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning

DGX agent

arXiv:2508.19679v2 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have enabled mobile agents to perceive and interact with real-world mobile environments based on hu

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training

DGX agent

arXiv:2506.20332v4 Announce Type: replace Abstract: Vision-language model-based mobile agents have gained the ability to understand complex instructions and mobile screenshots, benefiting from reinfor

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Scalable Agentic Reasoning for Designing Biologics Targeting Intrinsically Disordered Proteins

DGX agent

arXiv:2512.15930v2 Announce Type: replace-cross Abstract: Intrinsically disordered proteins (IDPs) represent crucial therapeutic targets due to their significant role in disease -- approximately 80% o

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

DGX agent

arXiv:2509.13021v2 Announce Type: replace-cross Abstract: This work introduces xOffense, an AI-driven, multi-agent penetration testing framework that shifts the process from labor-intensive, expert-dr

model-releasesarxiv-cs-ai
28 Apr 2026
Agents

A Probabilistic Framework for Hierarchical Goal Recognition

DGX agent

arXiv:2604.22256v1 Announce Type: cross Abstract: Goal recognition aims to infer an agent's goal from observations of its behaviour. In realistic settings, recognition can benefit from exploiting hier

agentsarxiv-cs-ai
27 Apr 2026
Agents

Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning

DGX agent

arXiv:2604.22191v1 Announce Type: cross Abstract: In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a

agentsarxiv-cs-cl
27 Apr 2026
Research

Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives

DGX agent

arXiv:2603.26041v3 Announce Type: replace Abstract: In recent years, GUI visual agents built upon Multimodal Large Language Models (MLLMs) have demonstrated strong potential in navigation tasks. Howev

researcharxiv-cs-cv
27 Apr 2026
Agents

SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning

DGX agent

arXiv:2604.22558v1 Announce Type: cross Abstract: As Multimodal Large Language Models (MLLMs) mature, GUI agents are evolving from static interactions to complex navigation. While Reinforcement Learni

agentsarxiv-cs-ai
27 Apr 2026
Agents

Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework

DGX agent

arXiv:2604.21090v1 Announce Type: cross Abstract: AI governance programmes increasingly rely on natural language prompts to constrain and direct AI agent behaviour. These prompts function as executabl

agentsarxiv-cs-ai
24 Apr 2026
Safety

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration

DGX agent

arXiv:2604.19856v1 Announce Type: cross Abstract: Large Language Models (LLMs) show promise for generating Register-Transfer Level (RTL) code from natural language specifications, but single-shot gene

safetyarxiv-cs-ai
23 Apr 2026
Safety

Environmental Understanding Vision-Language Model for Embodied Agent

DGX agent

arXiv:2604.19839v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these a

safetyarxiv-cs-ai
23 Apr 2026
Safety

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards

DGX agent

arXiv:2604.20486v1 Announce Type: new Abstract: Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outc

safetyarxiv-cs-cv
23 Apr 2026
Model Releases

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps

DGX agent

arXiv:2604.19533v1 Announce Type: cross Abstract: We introduce the Cyber Defense Benchmark, a benchmark for measuring how well large language model (LLM) agents perform the core SOC analyst task of th

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

SAGE-32B: Agentic Reasoning via Iterative Distillation

DGX agent

arXiv:2601.04237v2 Announce Type: replace Abstract: We demonstrate SAGE-32B, a 32 billion parameter language model that focuses on agentic reasoning and long range planning tasks. Unlike chat models t

model-releasesarxiv-cs-ai
22 Apr 2026
Agents

Generative midtended cognition and Artificial Intelligence. Thinging with thinging things

DGX agent

arXiv:2411.06812v2 Announce Type: replace-cross Abstract: This paper introduces the concept of ``generative midtended cognition'', exploring the integration of generative AI with human cognition. The

agentsarxiv-cs-lg
21 Apr 2026
Safety

OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

DGX agent

arXiv:2511.21064v2 Announce Type: replace-cross Abstract: Open-Vocabulary Object Detection (OVOD) aims to enable detectors to generalize across categories by leveraging semantic information. Although

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

DGX agent

arXiv:2604.15715v1 Announce Type: cross Abstract: The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows

model-releasesarxiv-cs-ai
20 Apr 2026
Agents

InfoChess: A Game of Adversarial Inference and a Laboratory for Quantifiable Information Control

DGX agent

arXiv:2604.15373v1 Announce Type: cross Abstract: We propose InfoChess, a symmetric adversarial game that elevates competitive information acquisition to the primary objective. There is no piece captu

agentsarxiv-cs-ai
20 Apr 2026
Safety

Long-Term Memory for VLA-based Agents in Open-World Task Execution

DGX agent

arXiv:2604.15671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated significant potential for embodied decision-making; however, their application in complex chemical

safetyarxiv-cs-ro
20 Apr 2026
Model Releases

PolicyBank: Evolving Policy Understanding for LLM Agents

DGX agent

arXiv:2604.15505v1 Announce Type: cross Abstract: LLM agents operating under organizational policies must comply with authorization constraints typically specified in natural language. In practice, su

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

DGX agent

arXiv:2604.14228v1 Announce Type: cross Abstract: Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes

model-releasesarxiv-cs-cl
17 Apr 2026
Safety

Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization

DGX agent

arXiv:2604.14267v1 Announce Type: new Abstract: Search agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and long-tail information unavaila

safetyarxiv-cs-lg
17 Apr 2026
Model Releases

SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment

DGX agent

arXiv:2604.13630v1 Announce Type: cross Abstract: The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

VeruSAGE: A Study of Agent-Based Verification for Rust Systems

DGX agent

arXiv:2512.18436v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown impressive capability to understand and develop code. However, their capability to rigorously reason a

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

DGX agent

arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration

DGX agent

arXiv:2604.14116v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have empowered AI research agents to perform isolated scientific tasks, automating complex, real-world workflows, s

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

DGX agent

arXiv:2512.14234v2 Announce Type: replace Abstract: Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition

DGX agent

arXiv:2604.12735v1 Announce Type: new Abstract: LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this p

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents

DGX agent

arXiv:2604.12948v1 Announce Type: new Abstract: LLM agents with persistent memory store information as flat factual records, providing little context for temporal reasoning, change tracking, or cross-

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

DGX agent

arXiv:2604.12616v1 Announce Type: new Abstract: The rapid evolution of Vision-Language Models (VLMs) has catalyzed unprecedented capabilities in artificial intelligence; however, this continuous modal

safetyarxiv-cs-ai
15 Apr 2026
← Previous
1…8384858687…236
Next →