AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

Long-term Fairness with Selective Labels

DGX agent

arXiv:2605.22291v1 Announce Type: new Abstract: Long-term fairness algorithms aim to satisfy fairness beyond static and short-term notions by accounting for the dynamics between decision-making polici

safetyarxiv-cs-lg
23 May 2026
Safety

MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2603.19310v3 Announce Type: replace Abstract: Reinforcement learning has emerged as a powerful paradigm for improving large language model (LLM) reasoning, where rollouts are sampled from the po

safetyarxiv-cs-lg
23 May 2026
Safety

MetaDNS: Enhancing Exploration in Discrete Neural Samplers via Well-Tempered Metadynamics

DGX agent

arXiv:2605.21722v1 Announce Type: cross Abstract: Sampling from discrete distributions with multiple modes and energy barriers is fundamental to machine learning and computational physics. Recent disc

safetyarxiv-cs-lg
23 May 2026
Safety

Objective-Induced Bias and Search Dynamics in Multiobjective Unsupervised Feature Selection

DGX agent

arXiv:2605.21561v1 Announce Type: new Abstract: Unsupervised feature selection is commonly formulated as a multiobjective optimisation problem that jointly optimises subset quality and subset size. Ye

safetyarxiv-cs-lg
23 May 2026
Safety

On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents

DGX agent

arXiv:2605.21763v1 Announce Type: new Abstract: We study risk-sensitive reinforcement learning in finite discounted MDPs, where a generative model of the MDP is assumed to be available. We consider a

safetyarxiv-cs-lg
23 May 2026
Safety

One-Way Policy Optimization for Self-Evolving LLMs

DGX agent

arXiv:2605.22156v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a promising paradigm for scaling reasoning capabilities of Large Language Models (LLMs)

safetyarxiv-cs-lg
23 May 2026
Safety

PEARL: Unbiased Percentile Estimation via Contrastive Learning for Industrial-Scale Livestream Recommendation

DGX agent

arXiv:2605.21752v1 Announce Type: new Abstract: Recommender systems trained on user interaction data are susceptible to behavioral intensity imbalance--a systematic distortion arising from heterogeneo

safetyarxiv-cs-lg
23 May 2026
Safety

Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation

DGX agent

arXiv:2605.22731v1 Announce Type: new Abstract: Large language model post-training methods such as supervised fine-tuning (SFT), reinforcement learning (RL), and distillation are often analyzed throug

safetyarxiv-cs-lg
23 May 2026
Safety

Proxy-Based Approximation of Shapley and Banzhaf Interactions

DGX agent

arXiv:2605.22738v1 Announce Type: new Abstract: Shapley and Banzhaf interactions capture the complex dynamics inherent in modern machine learning applications. However, current estimators for these hi

safetyarxiv-cs-lg
23 May 2026
Safety

[Re] FairDICE: A Fair Tradeoff in Multi-objective Offline RL

DGX agent

arXiv:2603.03454v2 Announce Type: replace Abstract: Offline Reinforcement Learning (RL) is an emerging field of RL in which policies are learned solely from demonstrations. Within offline RL, some env

safetyarxiv-cs-lg
23 May 2026
Safety

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

DGX agent

arXiv:2602.10894v2 Announce Type: replace Abstract: Two-player games such as board games have long been used as traditional benchmarks for reinforcement learning. This work revisits a policy optimizat

safetyarxiv-cs-lg
23 May 2026
Safety

SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression

DGX agent

arXiv:2511.04838v2 Announce Type: replace Abstract: Molecular property regression struggles with cases in chemically relevant target ranges that are underrepresented in datasets. Standard average erro

safetyarxiv-cs-lg
23 May 2026
Safety

Support-aware offline policy selection for advertising marketplaces

DGX agent

arXiv:2605.21736v1 Announce Type: cross Abstract: Logged advertising auctions make offline reserve-price evaluation attractive but risky. Replay tables can identify policies with large apparent yield

safetyarxiv-cs-lg
23 May 2026
Safety

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning

DGX agent

arXiv:2605.22263v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) is an emerging LLM post-training paradigm in which the model serves as its own teacher: conditioned on privileged inf

safetyarxiv-cs-lg
23 May 2026
Safety

Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning

DGX agent

arXiv:2605.22376v1 Announce Type: new Abstract: Cross-domain offline reinforcement learning (CDRL) aims to improve policy learning in a target domain by leveraging data collected from a source domain.

safetyarxiv-cs-lg
23 May 2026
Safety

The Attribution Impossibility: No Feature Ranking Is Faithful, Stable, and Complete Under Collinearity

DGX agent

arXiv:2605.21492v1 Announce Type: new Abstract: No feature ranking can be simultaneously faithful, stable, and complete when features are collinear. For collinear pairs, ranking reduces to a coin flip

safetyarxiv-cs-lg
23 May 2026
Safety

The Signal in the Noise: OOD Detection Through Goodness-of-Fit Testing in Factorised Latent Spaces

DGX agent

arXiv:2605.22496v1 Announce Type: new Abstract: Deep generative models offer a natural foundation for out-of-distribution (OOD) detection, yet prior work has shown that their assigned likelihoods are

safetyarxiv-cs-lg
23 May 2026
Safety

TreeDQN: Sample-Efficient Off-Policy Reinforcement Learning for Combinatorial Optimization

DGX agent

arXiv:2306.05905v2 Announce Type: replace Abstract: A convenient approach to optimally solving combinatorial optimization tasks is the Branch-and-Bound method. Its branching heuristic can be learned t

safetyarxiv-cs-lg
23 May 2026
Safety

Twice Sequential Monte Carlo for Tree Search

DGX agent

arXiv:2511.14220v3 Announce Type: replace Abstract: Model-based reinforcement learning (RL) methods that leverage search are responsible for many milestone breakthroughs in RL. Sequential Monte Carlo

safetyarxiv-cs-lg
23 May 2026
Safety

Uncertainty quantification for Markov chain induced martingales with application to temporal difference learning

DGX agent

arXiv:2502.13822v3 Announce Type: replace-cross Abstract: We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Ma

safetyarxiv-cs-lg
23 May 2026
Safety

What are the Right Symmetries for Formal Theorem Proving?

DGX agent

arXiv:2605.22257v1 Announce Type: new Abstract: Formal theorem provers based on large language models (LLMs) are highly sensitive to superficial variations in problem representation: semantically equi

safetyarxiv-cs-lg
23 May 2026
Safety

When to Switch, Not Just What: Transition Quality Prediction in Clash Royale

DGX agent

arXiv:2605.21868v1 Announce Type: new Abstract: In competitive games, players frequently switch strategies after losing streaks, yet our analysis of 926,334 match records from 34,619 Clash Royale play

safetyarxiv-cs-lg
23 May 2026
Safety

A KL-regularization Framework for Learning to Plan with Adaptive Priors

DGX agent

arXiv:2510.04280v2 Announce Type: replace-cross Abstract: Effective exploration remains a central challenge in model-based reinforcement learning (MBRL), particularly in high-dimensional continuous co

safetyarxiv-cs-ro
22 May 2026
Safety

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

DGX agent

arXiv:2605.22608v1 Announce Type: new Abstract: Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious

safetyarxiv-cs-cl
22 May 2026
Safety

Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction

DGX agent

arXiv:2605.21653v1 Announce Type: cross Abstract: AI text detectors amplify a pretrained typicality axis; they do not construct an AI-vs-human boundary. On raw encoders before any task supervision, pr

safetyarxiv-cs-cl
22 May 2026
Safety

Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations

DGX agent

arXiv:2605.22123v1 Announce Type: new Abstract: Designing reward functions that generalize beyond controlled laboratory settings remains a fundamental challenge in reinforcement learning for robotics.

safetyarxiv-cs-ro
22 May 2026
Safety

Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-Guided Refinement

DGX agent

arXiv:2605.22547v1 Announce Type: new Abstract: Deep learning has brought significant progress to medical image classification, yet most existing methods still rely on isolated visual evidence and can

safetyarxiv-cs-cv
22 May 2026
Safety

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

DGX agent

arXiv:2505.16416v3 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) is widely adopted in large language models, but when applied to vision-language models (VLMs) it couples text and i

safetyarxiv-cs-cv
22 May 2026
Safety

Closed-Loop Sim-to-Real Reinforcement Learning for Deformable Microfiber Shape Control

DGX agent

arXiv:2605.21688v1 Announce Type: new Abstract: Autonomous contact-based micromanipulation is challenging because surface and interfacial interactions at the microscale are difficult to model accurate

safetyarxiv-cs-ro
22 May 2026
Safety

Depth Augmented and FE Free 3D/2D Liver Registration for Laparoscopic Liver AR

DGX agent

arXiv:2602.17517v2 Announce Type: replace Abstract: Augmented reality (AR) guidance in laparoscopic liver surgery requires accurate registration of preoperative 3D models to intraoperative 2D video, b

safetyarxiv-cs-cv
22 May 2026
Safety

Discovering Implicit Large Language Model Alignment Objectives

DGX agent

arXiv:2602.15338v2 Announce Type: replace-cross Abstract: Large language model (LLM) alignment relies on complex reward signals that often obscure the specific behaviors being incentivized, creating c

safetyarxiv-cs-cl
22 May 2026
Safety

Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates

DGX agent

arXiv:2605.22061v1 Announce Type: new Abstract: Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core ch

safetyarxiv-cs-cv
22 May 2026
Safety

Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos

DGX agent

arXiv:2605.22066v1 Announce Type: new Abstract: Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and te

safetyarxiv-cs-cv
22 May 2026
Safety

Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention

DGX agent

arXiv:2605.21842v1 Announce Type: cross Abstract: Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intri

safetyarxiv-cs-cl
22 May 2026
Safety

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

DGX agent

arXiv:2605.22057v1 Announce Type: new Abstract: Enterprise routers assign queries to expert agents, yet deployed profiles stay static while agents evolve (prompts, tools, models), and developers rarel

safetyarxiv-cs-cl
22 May 2026
Safety

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

DGX agent

arXiv:2602.17186v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying

safetyarxiv-cs-cv
22 May 2026
Safety

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

DGX agent

arXiv:2605.22671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior

safetyarxiv-cs-cv
22 May 2026
Safety

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

DGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

safetyarxiv-cs-cv
22 May 2026
Safety

GLeVE: Graph-Guided Lesion Grounding with Proposal Verification in 3D CT

DGX agent

arXiv:2605.22619v1 Announce Type: new Abstract: Grounding radiology report descriptions to 3D CT volumes is essential for verifiable clinical interpretation, yet remains challenging due to the semanti

safetyarxiv-cs-cv
22 May 2026
Safety

Governance by Construction for Generalist Agents

DGX agent

arXiv:2605.20874v1 Announce Type: new Abstract: Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by constr

safetyarxiv-cs-ai
22 May 2026
Safety

Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting

DGX agent

arXiv:2605.22258v1 Announce Type: new Abstract: Large language models (LLMs) require robust toxicity evaluation beyond explicit wording. This setting remains underexplored in Chinese, where toxicity m

safetyarxiv-cs-cl
22 May 2026
Safety

Hierarchical Variational Policies for Reward-Guided Diffusion

DGX agent

arXiv:2605.21661v1 Announce Type: cross Abstract: Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We

safetyarxiv-cs-cv
22 May 2026
Safety

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance

DGX agent

arXiv:2605.22567v1 Announce Type: new Abstract: Reinforcement learning has proven effective for enhancing multi-step reasoning in large language models (LLMs), yet its benefits have not fully translat

safetyarxiv-cs-cl
22 May 2026
Safety

Learning Altruistic Collaboration in Heterogeneous Multi-Team Systems

DGX agent

arXiv:2605.21723v1 Announce Type: new Abstract: This paper studies heterogeneous multi-team collaboration through dynamic robot allocation, where robots are treated as transferable resources. Leveragi

safetyarxiv-cs-ro
22 May 2026
Safety

Learning to Configure Agentic AI Systems

DGX agent

arXiv:2602.11574v3 Announce Type: replace Abstract: Configuring LLM-based agent systems involves choosing workflows, tools, token budgets, and prompts from a large combinatorial design space, and is t

safetyarxiv-cs-ai
22 May 2026
Safety

Mapping Tomato Cropping Systems in California Using AlphaEarth Geospatial Embeddings and Deep Learning Analysis

DGX agent

arXiv:2605.21804v1 Announce Type: cross Abstract: Field-scale crop maps support supply-chain forecasting and policy, yet statewide crop identification still often depends on retrospective surveys or r

safetyarxiv-cs-cv
22 May 2026
Safety

Modeling Emotional Dynamics in Agent-to-Agent Interactions on Moltbook

DGX agent

arXiv:2605.20442v1 Announce Type: cross Abstract: Generative AI systems are increasingly deployed as interactive agents in online environments, such as a social network called Moltbook. In Moltbook, l

safetyarxiv-cs-ai
22 May 2026
Safety

Modeling Pathology-Like Behavioral Patterns in Language Models Through Behavioral Fine-Tuning

DGX agent

arXiv:2605.22356v1 Announce Type: new Abstract: Large language models are increasingly used as computational tools for modeling human-like behavior. We introduce a behavioral induction framework that

safetyarxiv-cs-cl
22 May 2026
← Previous
1…165166167168169…260
Next →