AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,973 results
11 Jun 2026

The Impossibility of Eliciting Latent Knowledge

AgentsDGX agent

arXiv:2606.12268v1 Announce Type: new Abstract: Advanced AI systems have extensive knowledge of their environments; in fact, their knowledge may (far) exceed that of their developers or users. Consequ

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning

SafetyDGX agent

arXiv:2606.12372v1 Announce Type: cross Abstract: Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy imp

10 Jun 2026

AnomaMind: Agentic Time Series Anomaly Detection with Tool-Augmented Reasoning

Local AiDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.13807v2 Announce Type: replace Abstract: Time series anomaly detection is critical in many real-world applications, where effective solutions must localize anomalous regions and support rel

Causal Ensemble Agent: Hierarchical Causal Discovery with LLM-guided Expert Reweighting

SafetyDGX agent

arXiv:2606.10607v1 Announce Type: cross Abstract: Causal discovery aims to uncover causal structures from observational data, which is crucial for real-world decision-making. However, different causal

Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness

ResearchDGX agent

arXiv:2606.10413v1 Announce Type: new Abstract: Breakthroughs in large language models and multimodal generation technologies have propelled the digital reconstruction of human mental traits, emotiona

9 Jun 2026

Byzantine Cheap Talk: Adversarial Resilience and Topology Effects in LLM Coordination Games

AgentsDGX agent

arXiv:2606.07790v1 Announce Type: new Abstract: Multi-agent LLM systems increasingly rely on communication protocols for coordination, yet their robustness under adversarial and structural constraints

Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech

ToolsDGX agent

This paper benchmarks frontier automatic speech recognition (ASR) systems on their ability to handle code-switched speech, where bilingual speakers mix languages within a single conversation. The rese

Evaluate Clinical ASR Models Faster with Agent Skills and NVIDIA Nemotron Speech

Model ReleasesDGX agent

This article presents a clinical automatic speech recognition workflow for generating pronunciation-aware synthetic audio, reviewing clinical terms, and evaluating recognition quality using NVIDIA age

Filigran launches XTM One to automate threat exposure management with AI agents

Model ReleasesDGX agent

French cybersecurity company Filigran SAS today launched XTM One, an artificial intelligence orchestration layer that automates continuous threat exposure management workflows across its platform. XTM

FineGen: A VLM-based Multi-Agent Framework for Fine-Grained Image-Text Dataset Construction

Model ReleasesDGX agent

arXiv:2606.07645v1 Announce Type: cross Abstract: The scarcity of hard negative samples in current vision-language datasets significantly hinders fine-grained perception. To address this, we propose F

HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents

Model ReleasesDGX agent

arXiv:2606.09738v1 Announce Type: new Abstract: Text-driven indoor scene generation and editing require an intermediate representation that language models can both produce and revise. Existing LLM-ba

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

SafetyDGX agent

arXiv:2606.08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes. Creations by

In-Context Reinforcement Learning via Communicative World Models

AgentsDGX agent

arXiv:2508.06659v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) agents often struggle to generalize to new tasks and contexts without updating their parameters, mainly because th

MAGIS: Evidence-Based Multi-Agent Reasoning for Interpretable Strabismus Clinical Decision-Making

Model ReleasesDGX agent

arXiv:2606.09249v1 Announce Type: new Abstract: Strabismus is a common ocular disorder that requires fine-grained subtype diagnosis for individualized treatment planning. However, existing deep learni

8 Jun 2026

Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning

SafetyDGX agent

arXiv:2511.02748v2 Announce Type: replace-cross Abstract: We argue that sixth-generation (6G) intelligence is not fluent token prediction but the capacity to imagine and choose -- to simulate future s

CRAFT: Coaching Reinforcement Learning Autonomously using Foundation Models for Multi-Robot Coordination Tasks

AgentsDGX agent

arXiv:2509.14380v3 Announce Type: replace Abstract: Multi-Agent Reinforcement Learning (MARL) provides a powerful framework for learning coordination in multi-agent systems. However, applying MARL to

tl;dr: we aren’t close to RSI, regardless of the hints IPO-bound Anthropic tried to drop last week.

AgentsDGX agent

tl;dr: we aren’t close to RSI, regardless of the hints IPO-bound Anthropic tried to drop last week. This paper tests whether today’s AI agents can build better AI agents without human design help. i.e

What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media

Model ReleasesDGX agent

arXiv:2606.06784v1 Announce Type: cross Abstract: Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulativ

6 Jun 2026

Evaluating Agentic Configuration Repair for Computer Networks

Model ReleasesDGX agent

arXiv:2606.06212v1 Announce Type: new Abstract: Misconfigurations in computer networks remain a major source of critical Internet outages. Research is turning to Large Language Models (LLMs) to automa

How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

Model ReleasesDGX agent

arXiv:2606.05256v1 Announce Type: new Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView. The intervention, conducted by unknown,

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows

Model ReleasesDGX agent

arXiv:2605.12376v2 Announce Type: replace Abstract: Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines

5 Jun 2026

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

Model ReleasesDGX agent

arXiv:2606.05557v1 Announce Type: new Abstract: A situated query like 'where is Lin Wei?' often encodes more than its literal content: the user may also want to know whether Lin Wei is free, in a good

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

SafetyDGX agent

arXiv:2606.06493v1 Announce Type: new Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

SafetyDGX agent

arXiv:2606.05743v1 Announce Type: cross Abstract: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifi

ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.05836v1 Announce Type: new Abstract: Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world d

The Meta hack shows there’s more to AI security than Mythos

AgentsDGX agent

On June 5, 404 Media reported that attackers had been using Meta’s AI customer support agent to steal Instagram accounts. Their approach was simple: They asked the agent to link the accounts to email

4 Jun 2026

Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal

AgentsDGX agent

arXiv:2606.04223v1 Announce Type: new Abstract: Multi-agent systems are commonly designed to reduce disagreement through voting, consensus protocols, debate, or fault-tolerant aggregation. We argue th

HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite

Model ReleasesDGX agent

arXiv:2606.04126v1 Announce Type: cross Abstract: We introduce HighTide, an evolving AI-assisted benchmark suite. Specifically, the contributions are: (i) a diverse open-source suite spanning multiple

Impostor: An Agent-Curated Benchmark for Realistic AIGC Manipulation Localization

Model ReleasesDGX agent

arXiv:2606.04545v1 Announce Type: new Abstract: Recent advances in generative image editing have improved the realism and controllability of localized image manipulation, raising new challenges for im

LifeSide: Benchmarking Agents as Lifelong Digital Companions

Model ReleasesDGX agent

arXiv:2606.04660v1 Announce Type: new Abstract: Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries

Neetyabhas: A Framework for Uncertainty-Aware Public Policy Optimization in Rational Agent-Based Models

SafetyDGX agent

arXiv:2606.04562v1 Announce Type: new Abstract: Purpose The WHO's COVID-19 non-pharmaceutical interventions (e.g., lockdowns, vaccinations) effectively curb transmission but impose heavy economic stra

Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

SafetyDGX agent

arXiv:2510.13704v2 Announce Type: replace-cross Abstract: Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment paralleliz

Sometimes you need to start over. But that decision is hard. @llama_index had to make that call: they built one of the most popular AI frame…

AgentsDGX agent

Sometimes you need to start over. But that decision is hard. @llama_index had to make that call: they built one of the most popular AI frameworks in the world, but saw the agent harness and frontier l

Towards Efficient and Evidence-grounded Mobility Prediction with LLM-Driven Agent

Model ReleasesDGX agent

arXiv:2606.05130v1 Announce Type: cross Abstract: Individual-level mobility prediction is central to urban simulation, transportation planning, and policy analysis. Supervised sequence models achieve

3 Jun 2026

AI Agents Enable Adaptive Computer Worms

SafetyDGX agent

arXiv:2606.03811v1 Announce Type: cross Abstract: A computer worm is malware that spreads on a network by replicating itself from one machine to another. Traditional worms, like WannaCry, exploited pr

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

Model ReleasesDGX agent

arXiv:2606.03829v1 Announce Type: new Abstract: Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and a

Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

AgentsDGX agent

arXiv:2510.23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

Model ReleasesDGX agent

arXiv:2606.03168v1 Announce Type: new Abstract: While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Model ReleasesDGX agent

arXiv:2606.03303v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages

Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles

HardwareDGX agent

arXiv:2505.08222v3 Announce Type: replace-cross Abstract: Autonomous vehicles (AVs) offer a cost-effective solution for scientific missions such as underwater tracking. Reinforcement learning (RL) has

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

Model ReleasesDGX agent

arXiv:2606.03980v1 Announce Type: cross Abstract: Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) p

This SkillOpt paper from Microsoft is a must-read! (bookmark it) I was a bit skeptical of the results reported in the paper when I shared it…

AgentsDGX agent

This SkillOpt paper from Microsoft is a must-read! (bookmark it) I was a bit skeptical of the results reported in the paper when I shared it a few days ago. However, I managed to integrate it into my

2 Jun 2026

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

Model ReleasesDGX agent

arXiv:2606.01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neur

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.02459v1 Announce Type: new Abstract: Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is d

AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

Model ReleasesDGX agent

arXiv:2606.02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical f

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

Model ReleasesDGX agent

arXiv:2602.18008v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on s

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

Model ReleasesDGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

Model ReleasesDGX agent

arXiv:2606.02215v1 Announce Type: new Abstract: Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social p

ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents

Model ReleasesDGX agent

arXiv:2606.02568v1 Announce Type: new Abstract: Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to

Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts

SafetyDGX agent

arXiv:2606.01441v1 Announce Type: new Abstract: Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve in

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

Model ReleasesDGX agent

arXiv:2606.01434v1 Announce Type: new Abstract: Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each ci

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

AgentsDGX agent

arXiv:2606.01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a m

Four insights you might have missed from theCUBE’s coverage of KB4-CON

AgentsDGX agent

Artificial intelligence agents are turning agent risk management into a frontline security priority. As digital workers begin touching email, financial systems, collaboration tools and business workfl

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs

Local AiDGX agent

arXiv:2606.00104v1 Announce Type: cross Abstract: Foundation models are increasingly used to drive autonomous systems, yet existing approaches either keep the model in a tight control loop, raising la

SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

Model ReleasesDGX agent

arXiv:2606.01912v1 Announce Type: new Abstract: Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferen

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

Model ReleasesDGX agent

arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existi

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

Model ReleasesDGX agent

arXiv:2606.01276v1 Announce Type: new Abstract: Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs)

1 Jun 2026

CoMem: Context Management with A Decoupled Long-Context Model

AgentsDGX agent

arXiv:2605.30842v1 Announce Type: new Abstract: Context management enables agentic models to solve long-horizon tasks through iterative summarization of previous interaction histories. However, this p

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

SafetyDGX agent

arXiv:2605.31354v1 Announce Type: new Abstract: Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

Model ReleasesDGX agent

arXiv:2605.31148v1 Announce Type: cross Abstract: Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into ac

← Previous
1…145146147148149…300
Next →