AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
25 Jun 2026

Towards a Bathroom-Centered Human-Building Digital Twin Framework for Indoor Safety Analysis

SafetyDGX agent

arXiv:2606.23292v2 Announce Type: replace-cross Abstract: Bathroom use is a critical safety challenge for older adults because wet surfaces, constrained layouts, limited support, and frequent posture

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

SafetyDGX agent

arXiv:2606.25034v1 Announce Type: new Abstract: General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversaria

23 Jun 2026

Safe-SAGE: Social-Semantic Adaptive Guidance for Safe Engagement through Laplace-Modulated Poisson Safety Functions

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2603.05497v3 Announce Type: replace Abstract: Traditional safety-critical control methods, such as control barrier functions, suffer from semantic blindness, exhibiting the same behavior around

4 Jun 2026

Listening to the Workforce: Measuring Construction Worker Safety Attitudes from Social Media Discourse Using LLMs

SafetyDGX agent

arXiv:2606.04450v1 Announce Type: new Abstract: Worker safety attitudes are key determinants of whether protective practices are applied or bypassed on construction sites. Yet measuring them at scale

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

SafetyDGX agent

arXiv:2606.04778v1 Announce Type: new Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent

28 May 2026

Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections

SafetyDGX agent

arXiv:2605.27418v1 Announce Type: cross Abstract: The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transp

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

Model ReleasesDGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection

SafetyDGX agent

arXiv:2605.28030v1 Announce Type: cross Abstract: Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversa

26 May 2026

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

SafetyDGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

25 May 2026

Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays

SafetyDGX agent

arXiv:2605.23351v1 Announce Type: new Abstract: We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while

BarrierSteer: LLM Safety via Learning Barrier Steering

SafetyDGX agent

arXiv:2602.20102v2 Announce Type: replace-cross Abstract: Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe co

19 May 2026

Generating Realistic Safety-Critical Scenarios for Vehicle-Pedestrian Interactions

SafetyDGX agent

arXiv:2605.17229v1 Announce Type: new Abstract: Automated driving system deployment requires rigorous validation across safety-critical vehicle-pedestrian interactions, yet real-world datasets rarely

Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents

SafetyDGX agent

arXiv:2605.17830v1 Announce Type: new Abstract: Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under

SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation

SafetyDGX agent

arXiv:2601.18442v3 Announce Type: replace Abstract: Autonomous Vehicle (AV) requires rigorous testing in safety-critical scenarios for safety validation, yet its validation is hindered by the high cos

15 May 2026

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability

SafetyDGX agent

arXiv:2605.14246v1 Announce Type: cross Abstract: Many safety-critical control problems are modeled as risk-sensitive partially observable Markov decision processes, where the controller must make dec

Selective Safety Steering via Value-Filtered Decoding

SafetyDGX agent

arXiv:2605.14746v1 Announce Type: new Abstract: While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of wor

Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

SafetyDGX agent

arXiv:2605.14174v1 Announce Type: new Abstract: Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet

12 May 2026

Internalizing Safety Understanding in Large Reasoning Models via Verification

SafetyDGX agent

arXiv:2605.08930v1 Announce Type: new Abstract: While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment para

Shields to Guarantee Probabilistic Safety in MDPs

SafetyDGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

Safety-Critical LiDAR-Inertial Odometry with On-Manifold Deterministic Protection Level

SafetyDGX agent

arXiv:2605.09383v1 Announce Type: new Abstract: In safety-critical scenarios, the protection level of the autonomous navigation system is crucial for enabling mobile robots to perform safe tasks. Howe

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

SafetyDGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

7 May 2026

A Closed-Form Dual-Barrier CBF Safety Filter for Holonomic Robots on Incrementally Built Occupancy Grid Maps

SafetyDGX agent

arXiv:2605.05182v1 Announce Type: new Abstract: We present a dual-barrier control barrier function (CBF) safety filter for real-time, safety-critical velocity control of holonomic robots operating in

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

Model ReleasesDGX agent

arXiv:2605.04572v1 Announce Type: cross Abstract: Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors l

6 May 2026

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

SafetyDGX agent

arXiv:2605.01147v1 Announce Type: new Abstract: As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety propertie

30 Apr 2026

Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics

SafetyDGX agent

arXiv:2604.26836v1 Announce Type: new Abstract: Predictive safety filters (PSFs) leverage model predictive control to enforce constraint satisfaction during deep reinforcement learning (RL) exploratio

Unifying Runtime Monitoring Approaches for Safety-Critical Machine Learning: Application to Vision-Based Landing

SafetyDGX agent

arXiv:2604.26411v1 Announce Type: new Abstract: Runtime monitoring is essential to ensure the safety of ML applications in safety-critical domains. However, current research is fragmented, with indepe

28 Apr 2026

LLM-Augmented Traffic Signal Control with LSTM-Based Traffic State Prediction and Safety-Constrained Decision Support

SafetyDGX agent

arXiv:2604.23902v1 Announce Type: new Abstract: Traffic signal control is a critical task in intelligent transportation systems, yet conventional fixed-time and rule-based methods often struggle to ad

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

SafetyDGX agent

arXiv:2604.24348v1 Announce Type: new Abstract: The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS

TSAssistant: A Human-in-the-Loop Agentic Framework for Automated Target Safety Assessment

SafetyDGX agent

arXiv:2604.23938v1 Announce Type: new Abstract: Target Safety Assessment (TSA) requires systematic integration of heterogeneous evidence, including genetic, transcriptomic, target homology, pharmacolo

21 Apr 2026

Cat-DPO: Category-Adaptive Safety Alignment

SafetyDGX agent

arXiv:2604.17299v1 Announce Type: new Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusin

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

SafetyDGX agent

arXiv:2604.17730v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging d

Guardrails in Logit Space: Safety Token Regularization for LLM Alignment

Model ReleasesDGX agent

arXiv:2604.17210v1 Announce Type: new Abstract: Fine-tuning well-aligned large language models (LLMs) on new domains often degrades their safety alignment, even when using benign datasets. Existing sa

20 Apr 2026

Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4

SafetyDGX agent

This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat

Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility

SafetyDGX agent

arXiv:2604.15579v1 Announce Type: cross Abstract: AI agents that interact with their environments through tools enable powerful applications, but in high-stakes business settings, unintended actions c

16 Apr 2026

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

SafetyDGX agent

arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req

5 Aug 2026

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

Model ReleasesDGX agent

arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations

31 Jul 2026

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

SafetyDGX agent

arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original,

30 Jul 2026

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

SafetyDGX agent

arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk

28 Jul 2026

On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

SafetyDGX agent

arXiv:2607.23365v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education.

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot…

SafetyDGX agent

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot of it is about people that just really, even if it's slight

23 Jul 2026

Learning Personalized Safety Interventions for Haptic Human-Robot Shared Control

SafetyDGX agent

arXiv:2607.19534v1 Announce Type: new Abstract: Haptic feedback provides an implicit channel for communicating safety intentions during human-robot shared control. Existing haptic guidance systems typ

15 Jul 2026

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

Model ReleasesDGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

9 Jul 2026

PB-OEL: A Performance-Bounded Online Ensemble Learning Framework With Mixed Feedback for Real-Time Safety Assessment

SafetyDGX agent

arXiv:2503.15581v2 Announce Type: replace Abstract: Real-time safety assessment is critical for ensuring the reliable operation of complex dynamic systems. However, obtaining full safety labels in rea

8 Jul 2026

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

SafetyDGX agent

arXiv:2607.05407v1 Announce Type: cross Abstract: Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child

2 Jul 2026

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

SafetyDGX agent

arXiv:2607.00402v1 Announce Type: cross Abstract: Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent metho

30 Jun 2026

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

Model ReleasesDGX agent

arXiv:2606.30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while th

26 Jun 2026

Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes

SafetyDGX agent

arXiv:2606.27210v1 Announce Type: new Abstract: We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS

5 Jun 2026

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning

SafetyDGX agent

arXiv:2606.05523v1 Announce Type: new Abstract: Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can byp

2 Jun 2026

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

SafetyDGX agent

arXiv:2606.02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate t

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization

SafetyDGX agent

arXiv:2510.09330v3 Announce Type: replace Abstract: Ensuring that large language models (LLMs) comply with safety requirements is a central challenge in AI deployment. Existing alignment approaches pr

1 Jun 2026

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

SafetyDGX agent

arXiv:2605.31073v1 Announce Type: new Abstract: Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do

23 May 2026

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

SafetyDGX agent

arXiv:2605.21834v1 Announce Type: new Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Con

4 May 2026

Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning

SafetyDGX agent

arXiv:2605.00667v1 Announce Type: new Abstract: Safety is a primary challenge in real-world reinforcement learning (RL). Formulating safety requirements as state-wise constraints has become a prominen

1 May 2026

Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under Uncertainty

SafetyDGX agent

arXiv:2604.27193v1 Announce Type: new Abstract: Automatic Emergency Braking (AEB) systems represent a safety-critical national interest, with the National Highway Traffic Safety Administration (NHTSA)

Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems

SafetyDGX agent

arXiv:2604.27358v1 Announce Type: new Abstract: As large language model (LLM) agents are deployed in high-stakes environments, the question of how safely to delegate subtasks to specialized sub-agents

24 Apr 2026

Intent Laundering: AI Safety Datasets Are Not What They Seem

Model ReleasesDGX agent

arXiv:2602.16729v3 Announce Type: replace-cross Abstract: We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isol

SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging

SafetyDGX agent

arXiv:2503.17239v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) is a common practice to adapt generalist models to specialized domains. However, recent studies show

Secure LLM Fine-Tuning via Safety-Aware Probing

Model ReleasesDGX agent

arXiv:2505.16737v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable success across many applications, but their ability to generate harmful content raises s

22 Apr 2026

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

SafetyDGX agent

arXiv:2604.19559v1 Announce Type: new Abstract: Construction workers are highly vulnerable to heat stress, yet tools that translate real-time physiological data into actionable safety intelligence rem

Safety-Critical Contextual Control via Online Riemannian Optimization with World Models

SafetyDGX agent

arXiv:2604.19639v1 Announce Type: cross Abstract: Modern world models are becoming too complex to admit explicit dynamical descriptions. We study safety-critical contextual control, where a Planner mu

← Previous
12345…238
Next →