AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Agents

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

DGX agent

arXiv:2408.03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. Thi

agentsarxiv-cs-ai
28 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA

DGX agent

arXiv:2607.14561v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong reasoning performance, but their tendency to hallucinate limits their reliability in knowledge

agentsarxiv-cs-cl
28 Jul 2026
Model Releases

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

DGX agent

arXiv:2607.22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments

DGX agent

arXiv:2607.22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental refer

model-releasesarxiv-cs-lg
27 Jul 2026
Agents

Active Trust Management for Successful Human-Robot Teaming: Moving from a Trust Repair to a Trust Satisficing Perspective

DGX agent

arXiv:2607.13595v1 Announce Type: new Abstract: Integrating mobile robots into human teams promises significant capability improvements for tasks such as searching hazardous environments. Unlike exist

agentsarxiv-cs-ro
16 Jul 2026
Safety

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

DGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

safetyarxiv-cs-ai
16 Jul 2026
Safety

CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform

DGX agent

arXiv:2607.12086v1 Announce Type: new Abstract: Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validat

safetyarxiv-cs-cl
15 Jul 2026
Model Releases

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

DGX agent

arXiv:2607.12216v1 Announce Type: cross Abstract: Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instr

model-releasesarxiv-cs-ai
15 Jul 2026
Safety

LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis

DGX agent

arXiv:2606.16149v2 Announce Type: replace Abstract: Rare disease diagnosis involves interpreting clinical and genetic findings through complex diagnostic reasoning. We investigated whether this reason

safetyarxiv-cs-ai
10 Jul 2026
Model Releases

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

DGX agent

arXiv:2607.07696v1 Announce Type: cross Abstract: Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the datab

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

DGX agent

arXiv:2607.07459v1 Announce Type: cross Abstract: We present EmbodiedGen V2, a generative 3D world engine for building executable sim-ready environments for embodied intelligence. Sim-ready 3D asset g

safetyarxiv-cs-cv
9 Jul 2026
Agents

RLVP: Penalize the Path, Reward the Outcome

DGX agent

arXiv:2607.07435v1 Announce Type: cross Abstract: Agents acting on our behalf in the real world (e.g. placing phone calls) must learn online from costly, often irreversible interactions rather than ch

agentsarxiv-cs-ai
9 Jul 2026
Model Releases

AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis

DGX agent

arXiv:2607.03656v1 Announce Type: cross Abstract: Large Language Models are increasingly used to turn natural-language requirements into code. In access control, that shortcut is dangerous: a generate

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

DGX agent

arXiv:2607.02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the c

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

DGX agent

arXiv:2510.11098v5 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks r

model-releasesarxiv-cs-cl
7 Jul 2026
Safety

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows

DGX agent

arXiv:2607.01465v1 Announce Type: new Abstract: Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means

safetyarxiv-cs-ai
3 Jul 2026
Model Releases

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

DGX agent

arXiv:2607.01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently. As these workflows become repeated

model-releasesarxiv-cs-ai
3 Jul 2026
Safety

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

DGX agent

arXiv:2607.01557v1 Announce Type: cross Abstract: Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns require tailored

safetyarxiv-cs-ai
3 Jul 2026
Safety

Episodic-to-Semantic Consolidation Without Identity Drift

DGX agent

arXiv:2607.01988v1 Announce Type: new Abstract: Long-running adaptive intelligent agents face a structural tension between knowledge consolidation and information integrity. Memory consolidation is co

safetyarxiv-cs-ai
3 Jul 2026
Agents

SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces

DGX agent

arXiv:2607.02345v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents increasingly automate software engineering tasks through reusable skills, natural-language instruction documen

agentsarxiv-cs-ai
3 Jul 2026
Tutorials

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments

DGX agent

arXiv:2607.01470v1 Announce Type: new Abstract: Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for

tutorialsarxiv-cs-ai
3 Jul 2026
Model Releases

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

DGX agent

arXiv:2607.00304v1 Announce Type: cross Abstract: The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strat

model-releasesarxiv-cs-ai
2 Jul 2026
Safety

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

DGX agent

arXiv:2510.24636v3 Announce Type: replace Abstract: Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both traini

safetyarxiv-cs-cl
2 Jul 2026
Model Releases

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

DGX agent

arXiv:2606.31167v1 Announce Type: cross Abstract: VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current s

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

DGX agent

arXiv:2606.30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control

safetyarxiv-cs-ai
1 Jul 2026
Model Releases

A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

DGX agent

arXiv:2606.29719v1 Announce Type: cross Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it.

model-releasesarxiv-cs-cl
30 Jun 2026
Safety

Agentic Safety is an Epistemic Property, Not a Behavioral One

DGX agent

arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods

safetyarxiv-cs-ai
30 Jun 2026
Agents

Improved Multi-Dimensional Forecasting for Swap Regret

DGX agent

arXiv:2606.29533v1 Announce Type: cross Abstract: We study the problem of forecasting for an arbitrary number of downstream agents with unknown objectives, each of whom best responds to the forecaster

agentsarxiv-cs-lg
30 Jun 2026
Agents

Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts

DGX agent

arXiv:2606.29279v1 Announce Type: cross Abstract: LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into store

agentsarxiv-cs-ai
30 Jun 2026
Agents

On the Necessity of a Liquid Substrate for Mesh Intelligence

DGX agent

arXiv:2606.28413v1 Announce Type: cross Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain. Its competence rests on each

agentsarxiv-cs-ai
30 Jun 2026
Safety

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

DGX agent

arXiv:2606.28710v1 Announce Type: new Abstract: We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that

safetyarxiv-cs-ai
30 Jun 2026
Safety

Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities

DGX agent

arXiv:2606.26933v1 Announce Type: cross Abstract: AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violations and effi

safetyarxiv-cs-ai
26 Jun 2026
Agents

SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward

DGX agent

arXiv:2606.25195v1 Announce Type: cross Abstract: The increasing use of AI systems for code generation raises a central security question: what can today's models and coding agents actually do to prod

agentsarxiv-cs-ai
25 Jun 2026
Model Releases

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

DGX agent

arXiv:2606.23937v1 Announce Type: cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test t

model-releasesarxiv-cs-ai
24 Jun 2026
Agents

Position: Correct Answer, Wrong Mechanism -- When AI Scientists Defend General Claims Their Own Data Contradicts

DGX agent

arXiv:2606.23175v1 Announce Type: new Abstract: AI scientist systems are described as tools, coauthors, or founders, but we evaluate them as if only the final answer matters. This position paper argue

agentsarxiv-cs-lg
23 Jun 2026
Safety

Fourier Features Let Agents Learn High Precision Policies with Imitation Learning

DGX agent

arXiv:2606.12334v1 Announce Type: new Abstract: High-precision robotic manipulation requires fine-grained spatial reasoning that is often difficult to achieve with RGB-only policies due to depth ambig

safetyarxiv-cs-lg
11 Jun 2026
Model Releases

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

DGX agent

arXiv:2407.20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities. Large Language Model (LLM), which exhibits powerful language und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Constructing coherent spatial memory in LLM agents through graph rectification

DGX agent

arXiv:2510.04195v2 Announce Type: replace Abstract: Given a map description through global traversal navigation instructions, an LLM can often infer the implicit spatial layout and answer user queries

model-releasesarxiv-cs-ai
10 Jun 2026
Agents

Regimes: An Auditable, Held-Out-Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph

DGX agent

arXiv:2606.10241v1 Announce Type: new Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlogg

agentsarxiv-cs-ai
10 Jun 2026
Applications

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents

DGX agent

arXiv:2606.10457v1 Announce Type: new Abstract: Decision rules that enterprise experts apply tacitly -- in auditing, compliance, and contract review -- can be systematically recovered and improved thr

applicationsarxiv-cs-ai
10 Jun 2026
Agents

On the Hardness of Optimal Motion on Trees

DGX agent

arXiv:2606.06686v1 Announce Type: new Abstract: This paper presents a simple framework that settles the complexity of Multi-Agent Path Finding (MAPF) on trees across standard objectives--distance, mak

agentsarxiv-cs-ro
8 Jun 2026
Research

PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation

DGX agent

arXiv:2606.05697v1 Announce Type: new Abstract: User interface (UI) and user experience (UX) evaluation is central to product development, yet reliable feedback still relies on recruiting human partic

researcharxiv-cs-ai
6 Jun 2026
Agents

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

DGX agent

arXiv:2606.06473v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE),

agentsarxiv-cs-cl
5 Jun 2026
Local Ai

AgenticDiffusion: Agentic Diffusion-based Path Planning for Vision-Based UAV Navigation

DGX agent

arXiv:2606.04111v1 Announce Type: cross Abstract: Indoor UAV navigation requires efficient exploration, scene understanding, and reliable trajectory execution under limited field-of-view observations.

local-aiarxiv-cs-ai
4 Jun 2026
Safety

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

DGX agent

arXiv:2606.04703v1 Announce Type: new Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward c

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

DGX agent

arXiv:2606.03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a functi

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Knowing Isn't Understanding: Re-grounding Generative Proactivity with Epistemic and Behavioral Insight

DGX agent

arXiv:2602.15259v2 Announce Type: replace-cross Abstract: Generative AI agents equate understanding with resolving explicit queries, an assumption that confines interaction to what users can articulat

agentsarxiv-cs-ai
2 Jun 2026
Agents

VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding

DGX agent

arXiv:2602.04094v2 Announce Type: replace Abstract: Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints an

agentsarxiv-cs-cv
2 Jun 2026
← Previous
1…117118119120121…236
Next →