AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Local Ai

Playing games with knowledge: AI-Induced delusions need game theoretic interventions

DGX agent

arXiv:2605.08409v1 Announce Type: new Abstract: Conversational AI has a fundamental flaw as a knowledge interface: sycophantic chatbots induce epistemic entrenchment and delusional belief spirals even

local-aiarxiv-cs-ai
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Position: AI Security Policy Should Target Systems, Not Models

DGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

model-releasesarxiv-cs-ai
12 May 2026
Safety

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias

DGX agent

arXiv:2605.08315v1 Announce Type: new Abstract: Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole

safetyarxiv-cs-lg
12 May 2026
Safety

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

DGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

safetyarxiv-cs-cl
12 May 2026
Model Releases

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind

DGX agent

arXiv:2603.26089v2 Announce Type: replace-cross Abstract: The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind

model-releasesarxiv-cs-ai
12 May 2026
Safety

Shields to Guarantee Probabilistic Safety in MDPs

DGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

safetyarxiv-cs-ai
12 May 2026
Model Releases

Statistical Model Checking of the Keynes+Schumpeter Model: A Transient Sensitivity Analysis of a Macroeconomic ABM

DGX agent

arXiv:2605.10447v1 Announce Type: cross Abstract: Agent-based models (ABMs) are increasingly used in macroeconomics, but their analysis still often relies on ad hoc Monte Carlo campaigns with heteroge

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Step Rejection Fine-Tuning: A Practical Distillation Recipe

DGX agent

arXiv:2605.10674v1 Announce Type: cross Abstract: Rejection Fine-Tuning (RFT) is a standard method for training LLM agents, where unsuccessful trajectories are discarded from the training set. In the

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

DGX agent

arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas

model-releasesarxiv-cs-cl
12 May 2026
Safety

TIE: Time Interval Encoding for Video Generation over Events

DGX agent

arXiv:2605.10543v1 Announce Type: new Abstract: Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which

safetyarxiv-cs-cv
12 May 2026
Model Releases

Towards Conversational Medical AI with Eyes, Ears and a Voice

DGX agent

arXiv:2605.09272v1 Announce Type: new Abstract: The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchange and interpretation of rich auditory and visual cues bet

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews

DGX agent

arXiv:2605.10171v1 Announce Type: cross Abstract: Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for A

model-releasesarxiv-cs-ai
12 May 2026
Applications

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models

DGX agent

arXiv:2605.07308v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have significantly advanced the capabilities of robotic agents in executing diverse tasks; however, they still face

applicationsarxiv-cs-ro
11 May 2026
Model Releases

Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning

DGX agent

arXiv:2605.07333v1 Announce Type: new Abstract: In-context reinforcement learning (ICRL) studies agents that, after pretraining, adapt to new tasks by conditioning on additional context without parame

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching

DGX agent

arXiv:2601.15884v2 Announce Type: replace Abstract: Contrast-enhanced imaging is central to oncologic diagnosis, but contrast agents can be contraindicated for many of the patients who need them most.

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought

DGX agent

arXiv:2605.07123v1 Announce Type: new Abstract: In-context reinforcement learning (ICRL) refers to the ability of RL agents to adapt to new tasks at inference time without parameter updates by conditi

model-releasesarxiv-cs-lg
11 May 2026
Safety

Decentralized Time-Varying Optimization for Streaming Data via Temporal Weighting

DGX agent

arXiv:2605.06971v1 Announce Type: cross Abstract: Classical optimization theory largely focuses on fixed objective functions, whereas many modern learning systems operate in dynamic environments where

safetyarxiv-cs-ai
11 May 2026
Model Releases

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

DGX agent

arXiv:2605.07323v1 Announce Type: new Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regres

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators

DGX agent

arXiv:2605.06997v1 Announce Type: new Abstract: Long chain-of-thought reasoning and agentic tool-calling produce traces spanning tens of thousands of tokens, yet Transformer KV caches grow linearly wi

model-releasesarxiv-cs-lg
11 May 2026
Safety

Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning

DGX agent

arXiv:2605.06156v2 Announce Type: replace-cross Abstract: Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture co

safetyarxiv-cs-ai
11 May 2026
Safety

Multi-Environment POMDPs with Finite-Horizon Objectives

DGX agent

arXiv:2605.07537v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are systems in which one agent interacts with a stochastic environment, and receives only partia

safetyarxiv-cs-ai
11 May 2026
Model Releases

Multi-Objective Constraint Inference using Inverse reinforcement learning

DGX agent

arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin

model-releasesarxiv-cs-ai
11 May 2026
Safety

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

DGX agent

arXiv:2512.23770v3 Announce Type: replace-cross Abstract: In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tas

safetyarxiv-cs-ai
11 May 2026
Safety

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

DGX agent

arXiv:2605.07447v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly and are increasingly deployed in real-world applications, especially with the rise of agent-based

safetyarxiv-cs-ai
11 May 2026
Safety

Distilling Bayesian Belief States into Language Models for Auditable Negotiation

DGX agent

arXiv:2605.04507v1 Announce Type: new Abstract: Negotiation agents must infer what their counterpart values, update those beliefs over dialogue turns, and choose actions under uncertainty. End-to-end

safetyarxiv-cs-cl
7 May 2026
Model Releases

A Benchmark for Interactive World Models with a Unified Action Generation Framework

DGX agent

arXiv:2605.03941v1 Announce Type: new Abstract: Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable env

model-releasesarxiv-cs-cv
6 May 2026
Research

A Sentence Relation-Based Approach to Sanitizing Malicious Instructions

DGX agent

arXiv:2605.01078v1 Announce Type: cross Abstract: Retrieval-augmented generation and tool-integrated LLM agents increasingly depend on external textual sources. This reliance broadens the available at

researcharxiv-cs-ai
6 May 2026
Safety

Intervention Complexity as a Canonical Reward and a Measure of Intelligence

DGX agent

arXiv:2605.02175v1 Announce Type: new Abstract: The Legg--Hutter universal intelligence measure provides a rigorous scalar assessment of general intelligence as expected reward across all computable e

safetyarxiv-cs-ai
6 May 2026
Research

MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing

DGX agent

arXiv:2605.02199v1 Announce Type: new Abstract: Long-term LLM agents must compress streams of past interactions into persistent memory before future queries are known. Existing evaluations usually mea

researcharxiv-cs-ai
6 May 2026
Local Ai

On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization

DGX agent

arXiv:2511.11362v2 Announce Type: replace-cross Abstract: On-device fine-tuning is a critical capability for edge AI systems, which must support adaptation to different agentic tasks under stringent m

local-aiarxiv-cs-cl
6 May 2026
Safety

Poly-EPO: Training Exploratory Reasoning Models

DGX agent

arXiv:2604.17654v3 Announce Type: replace Abstract: Exploration is a cornerstone of learning from experience: it enables agents to find solutions to complex problems, generalize to novel ones, and sca

safetyarxiv-cs-ai
6 May 2026
Model Releases

Safety and accuracy follow different scaling laws in clinical large language models

DGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

model-releasesarxiv-cs-cl
6 May 2026
Safety

The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence

DGX agent

arXiv:2408.12622v3 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet resea

safetyarxiv-cs-lg
6 May 2026
Model Releases

Towards Understanding Specification Gaming in Reasoning Models

DGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

model-releasesarxiv-cs-ai
6 May 2026
Tutorials

ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable

DGX agent

arXiv:2604.25224v2 Announce Type: replace Abstract: LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a del

tutorialsarxiv-cs-ai
6 May 2026
Safety

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

DGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

safetyarxiv-cs-cv
5 May 2026
Model Releases

Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance

DGX agent

arXiv:2605.01368v1 Announce Type: new Abstract: Human-robot interaction (HRI) has long studied how agents and people coordinate to achieve shared goals. In this work, we formalize and benchmark the no

model-releasesarxiv-cs-ro
5 May 2026
Safety

Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs

DGX agent

arXiv:2605.01242v1 Announce Type: new Abstract: Reinforcement learning (RL) is a fundamental framework for sequential decision-making, in which an agent learns an optimal policy through interactions w

safetyarxiv-cs-lg
5 May 2026
Research

GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory

DGX agent

arXiv:2605.01688v1 Announce Type: new Abstract: Long-horizon conversational agents rely on memory systems with increasingly sophisticated retrieval mechanisms. However, retrieved fragments are typical

researcharxiv-cs-cl
5 May 2026
Local Ai

Hierarchical Federated Learning for Networked AI: From Communication Saving to Architecture-Aware Design

DGX agent

arXiv:2605.00931v1 Announce Type: new Abstract: Federated learning (FL) is fundamentally a distributed optimization problem executed by communicating agents with local data, local computation, and par

local-aiarxiv-cs-lg
5 May 2026
Research

HyMem: Hybrid Memory Architecture with Dynamic Retrieval Scheduling

DGX agent

arXiv:2602.13933v2 Announce Type: replace Abstract: Large language model (LLM) agents demonstrate strong performance in short-text contexts but often underperform in extended dialogues due to ineffici

researcharxiv-cs-ai
5 May 2026
Safety

Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare

DGX agent

arXiv:2605.01961v1 Announce Type: new Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents. However

safetyarxiv-cs-lg
5 May 2026
Model Releases

Robust volatility updates for Hierarchical Gaussian Filtering

DGX agent

arXiv:2605.00966v1 Announce Type: new Abstract: Hierarchical Gaussian Filtering (HGF) networks allow for efficient updating of posterior distributions (beliefs) about hidden states of an agent's envir

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't

DGX agent

arXiv:2605.01771v1 Announce Type: new Abstract: An auditor instructs an AI assistant: 'open each file individually using the Read tool -- no scripts, no agents.' The AI replies 'Yes' -- then issues a

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models

DGX agent

arXiv:2605.02363v1 Announce Type: new Abstract: Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two math

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

LLM-Oriented Information Retrieval: A Denoising-First Perspective

DGX agent

arXiv:2605.00505v1 Announce Type: cross Abstract: Modern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented g

model-releasesarxiv-cs-cl
4 May 2026
Applications

Scaling Federated Linear Contextual Bandits via Sketching

DGX agent

arXiv:2605.00500v1 Announce Type: new Abstract: In federated contextual linear bandits, high data dimensionality incurs prohibitive computation and communication costs: local agents perform O(d^3)-tim

applicationsarxiv-cs-lg
4 May 2026
Model Releases

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

DGX agent

arXiv:2602.14276v2 Announce Type: replace Abstract: Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain

model-releasesarxiv-cs-cv
4 May 2026
← Previous
1…181182183184185…233
Next →