AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,816 results
Safety

Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities

DGX agent

arXiv:2605.27388v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly utilized as proxies for computational social analysis; yet, their ability to faithfully represent the 't

safetyarxiv-cs-ai
28 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning

DGX agent

arXiv:2605.28552v1 Announce Type: new Abstract: As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle ty

safetyarxiv-cs-ai
28 May 2026
Safety

Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction

DGX agent

arXiv:2605.27878v1 Announce Type: new Abstract: Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and

safetyarxiv-cs-cl
28 May 2026
Safety

NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning

DGX agent

arXiv:2601.19947v2 Announce Type: replace-cross Abstract: Learning from Noisy Labels (LNL) remains a fundamental challenge in deep learning because real-world datasets often contain corrupted annotati

safetyarxiv-cs-ai
28 May 2026
Safety

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

DGX agent

arXiv:2603.01766v2 Announce Type: replace Abstract: Despite the rapid progress of vision-language-action (VLA) models, the prevailing practice of predicting action chunks as discrete waypoints remains

safetyarxiv-cs-ro
28 May 2026
Safety

No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees

DGX agent

arXiv:2603.08761v2 Announce Type: replace-cross Abstract: We argue that formal certification of AI alignment over open-ended or unbounded input domains is impossible under standard assumptions in comp

safetyarxiv-cs-lg
28 May 2026
Safety

No Safe Dose: How Training Data Drives Unsafe Image Generation

DGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

safetyarxiv-cs-cv
28 May 2026
Safety

Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think

DGX agent

arXiv:2605.28150v1 Announce Type: new Abstract: Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged

safetyarxiv-cs-lg
28 May 2026
Safety

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

DGX agent

arXiv:2604.18530v2 Announce Type: replace Abstract: Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet m

safetyarxiv-cs-ai
28 May 2026
Safety

@OpenAI Foundation just launched with a 25B commitment and equity in OpenAI valued at ~130B. Sounds like a philanthropy giant. But look cl…

DGX agent

@OpenAI Foundation just launched with a 25B commitment and equity in OpenAI valued at ~130B. Sounds like a philanthropy giant. But look closer: the big money goes to programs they control internally (

safetygary-marcus--x
28 May 2026
Safety

OpenAI’s Frontier Governance Framework

DGX agent

OpenAI's Frontier Governance Framework outlines the organization's approach to managing risks associated with advanced AI systems, including safety, security, and responsible deployment practices. The

safetyopenai
28 May 2026
Safety

Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions -- A Governance Framework for High-Stakes AI Systems

DGX agent

arXiv:2605.27827v1 Announce Type: new Abstract: AI governance frameworks increasingly emphasize fairness, transparency, accountability, and lifecycle risk management in high-stakes domains. However, m

safetyarxiv-cs-ai
28 May 2026
Safety

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

DGX agent

arXiv:2605.28675v1 Announce Type: new Abstract: Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are cos

safetyarxiv-cs-lg
28 May 2026
Safety

Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

DGX agent

arXiv:2605.28345v1 Announce Type: new Abstract: Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, an

safetyarxiv-cs-ai
28 May 2026
Safety

Position: Retire the 'Positive Backdoor' Label -- Secret Alignment Requires Strict and Systematic Evaluation

DGX agent

arXiv:2605.28597v1 Announce Type: cross Abstract: This position paper argues that the AI/ML community should stop overclaiming and retire the label 'positive backdoor,' and instead treat trigger-activ

safetyarxiv-cs-ai
28 May 2026
Safety

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

DGX agent

arXiv:2602.01745v2 Announce Type: replace-cross Abstract: Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dime

safetyarxiv-cs-ai
28 May 2026
Safety

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

DGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

safetyarxiv-cs-cl
28 May 2026
Safety

ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

DGX agent

arXiv:2605.28293v1 Announce Type: cross Abstract: Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinf

safetyarxiv-cs-ai
28 May 2026
Safety

Provably Guaranteed Polytopic Uncertainty Quantification for SLAM

DGX agent

arXiv:2605.28172v1 Announce Type: new Abstract: In safety-critical robotics applications, guaranteed and practical uncertainty quantification (UQ) in perception is vital. Many existing works either of

safetyarxiv-cs-ro
28 May 2026
Safety

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

DGX agent

arXiv:2602.15894v2 Announce Type: replace Abstract: In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, exist

safetyarxiv-cs-cl
28 May 2026
Safety

RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?

DGX agent

arXiv:2605.27436v1 Announce Type: cross Abstract: Multimodal alignment is critical for bridging the semantic gap in information retrieval. However, traditional pairwise strategies introduce a geometri

safetyarxiv-cs-ai
28 May 2026
Safety

Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization

DGX agent

arXiv:2605.27906v1 Announce Type: new Abstract: Multimodal Large Reasoning Models introduce the reasoning paradigm, demonstrating strong capabilities on complex vision-language tasks. However, they st

safetyarxiv-cs-ai
28 May 2026
Safety

Refining Multidimensional Video Reward Models via Disentangled Influence Functions

DGX agent

arXiv:2605.28203v1 Announce Type: new Abstract: As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various ax

safetyarxiv-cs-lg
28 May 2026
Safety

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

DGX agent

arXiv:2605.28553v1 Announce Type: new Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on

safetyarxiv-cs-ai
28 May 2026
Safety

Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning

DGX agent

arXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinforcement learning with large language models by lev

safetyarxiv-cs-ai
28 May 2026
Safety

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

DGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

safetyarxiv-cs-cl
28 May 2026
Safety

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

DGX agent

arXiv:2605.27996v1 Announce Type: new Abstract: Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto cor

safetyarxiv-cs-ai
28 May 2026
Safety

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

DGX agent

arXiv:2605.27834v1 Announce Type: new Abstract: We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in

safetyarxiv-cs-lg
28 May 2026
Safety

Right now, 20 British MPs are deciding which bills they'll introduce to Parliament. Sir Stephen Fry just asked them to bring forward our bil…

DGX agent

Right now, 20 British MPs are deciding which bills they'll introduce to Parliament. Sir Stephen Fry just asked them to bring forward our bill to ban superintelligence! Grateful to receive this strong

safetyconnor-leahy--x
28 May 2026
Safety

ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

DGX agent

arXiv:2509.13177v2 Announce Type: replace Abstract: Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their develop

safetyarxiv-cs-ro
28 May 2026
Safety

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

DGX agent

arXiv:2605.28014v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-

safetyarxiv-cs-cl
28 May 2026
Safety

Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models

DGX agent

arXiv:2605.28306v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have emerged as a dominant paradigm for efficient LLM scaling, yet adapting them to non-English downstream tasks remai

safetyarxiv-cs-ai
28 May 2026
Safety

SA4Depth: Consistent Pose-Depth Scale Alignment for Self-Supervised Monocular Depth Estimation

DGX agent

arXiv:2605.28477v1 Announce Type: new Abstract: Self-supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to

safetyarxiv-cs-cv
28 May 2026
Safety

SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

DGX agent

arXiv:2605.28338v1 Announce Type: new Abstract: Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance

safetyarxiv-cs-ai
28 May 2026
Safety

Safety-Critical Adaptive Impedance Control via Nonsmooth Control Barrier Functions under State and Input Constraints

DGX agent

arXiv:2605.28367v1 Announce Type: new Abstract: Safe physical interaction is critical for deploying robotic manipulators in human-robot interaction and contact-rich tasks, where uncertainty, external

safetyarxiv-cs-ro
28 May 2026
Safety

SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving

DGX agent

arXiv:2605.28583v1 Announce Type: cross Abstract: Ensuring both safety and efficiency in decision-making for autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement

safetyarxiv-cs-ai
28 May 2026
Safety

SCALE-COMM: Shared, Contrastively-Aligned Latent Embeddings for MARL Communication

DGX agent

arXiv:2605.27532v1 Announce Type: new Abstract: Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement

safetyarxiv-cs-ro
28 May 2026
Safety

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

DGX agent

arXiv:2605.27938v1 Announce Type: new Abstract: Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it

safetyarxiv-cs-cv
28 May 2026
Safety

Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity

DGX agent

arXiv:2605.27526v1 Announce Type: cross Abstract: We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regre

safetyarxiv-cs-lg
28 May 2026
Safety

Sense Representations Are Inducible Interfaces

DGX agent

arXiv:2605.28669v1 Announce Type: cross Abstract: Sense representations (explicit, per-token meaning decompositions) are useful for disambiguation, steering, and cross-lingual alignment, but existing

safetyarxiv-cs-ai
28 May 2026
Safety

Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning

DGX agent

arXiv:2605.27697v1 Announce Type: cross Abstract: Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing

safetyarxiv-cs-ai
28 May 2026
Safety

Singular Vectors of Attention Heads Align with Features

DGX agent

arXiv:2602.13524v2 Announce Type: replace-cross Abstract: Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the

safetyarxiv-cs-ai
28 May 2026
Safety

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

DGX agent

arXiv:2605.28791v1 Announce Type: cross Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense

safetyarxiv-cs-ai
28 May 2026
Safety

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

DGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

safetyarxiv-cs-ai
28 May 2026
Safety

Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams

DGX agent

arXiv:2605.27404v1 Announce Type: cross Abstract: The era of Big Science has long been defined by increasingly large and specialized research teams pushing the frontiers of knowledge. However, recent

safetyarxiv-cs-ai
28 May 2026
Safety

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

DGX agent

arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be che

safetyarxiv-cs-cl
28 May 2026
Safety

SPAR: Support-Preserving Action Rectification

DGX agent

arXiv:2605.27877v1 Announce Type: cross Abstract: Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression

safetyarxiv-cs-ai
28 May 2026
Safety

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection

DGX agent

arXiv:2605.28030v1 Announce Type: cross Abstract: Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversa

safetyarxiv-cs-ai
28 May 2026
← Previous
1…138139140141142…267
Next →