AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
61+ results
24 Apr 2026

Full-Body Dynamic Safety for Robot Manipulators: 3D Poisson Safety Functions for CBF-Based Safety Filters

SafetyDGX agent

arXiv:2604.21189v1 Announce Type: new Abstract: Collision avoidance for robotic manipulators requires enforcing full-body safety constraints in high-dimensional configuration spaces. Control Barrier F

10 Jul 2026

OpenAI's head of safety, Johannes Heidecke, is leaving as OpenAI integrates its research and safety teams; Mia Glaese will become VP of research and safety (Maxwell Zeff/Wired)

SafetyDGX agent

Maxwell Zeff / Wired: OpenAI's head of safety, Johannes Heidecke, is leaving as OpenAI integrates its research and safety teams; Mia Glaese will become VP of research and safety — Johannes Heidecke's

15 Apr 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Deep QP Safety Filter: Model-free Learning for Reachability-based Safety Filter

SafetyDGX agent

arXiv:2601.21297v2 Announce Type: replace Abstract: We introduce Deep QP Safety Filter, a fully data-driven safety layer for black-box dynamical systems. Our method learns a Quadratic-Program (QP) saf

3 Aug 2026

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Model ReleasesDGX agent

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, I

4 Jun 2026

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

SafetyDGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

Model ReleasesDGX agent

Nemotron 3.5 Content Safety is NVIDIA's multimodal safety solution designed for enterprise AI applications, offering customizable safeguards for both text and image inputs across different global cont

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Model ReleasesDGX agent

arXiv:2603.10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never

COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection

SafetyDGX agent

arXiv:2606.04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-valu

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

SafetyDGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

28 Apr 2026

Cooptimizing Safety and Performance Using Safety Value-Constrained Model Predictive Control

SafetyDGX agent

arXiv:2604.23863v1 Announce Type: new Abstract: Autonomous systems are increasingly deployed in real-world environments, where they must achieve high performance while maintaining safety under state a

Safety-aware Goal-oriented Semantic Sensing, Communication, and Control for Robotics

SafetyDGX agent

arXiv:2603.13502v2 Announce Type: replace Abstract: Wirelessly-connected robotic systems empower robots with real-time intelligence by leveraging remote computing resources for decision-making. Howeve

Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

SafetyDGX agent

arXiv:2604.23775v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of safety challenges, s

11 Aug 2026

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

Local AiDGX agent

arXiv:2608.09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial i

Beyond 'I Can't Help With That': How Child Safety Experts Evaluate AI Chatbot Safety

SafetyDGX agent

arXiv:2608.07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes s

10 Aug 2026

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

SafetyDGX agent

arXiv:2604.00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Mod

2 Jun 2026

Permissive Safety Through Trusted Inference: Verifiable Belief-Space Neural Safety Filters for Assured Interactive Robotics

Model ReleasesDGX agent

arXiv:2606.02562v1 Announce Type: cross Abstract: Autonomous robots that interact with people must make safe and efficient decisions under human-induced uncertainty, such as their preferences, goals,

Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

SafetyDGX agent

arXiv:2503.11832v5 Announce Type: replace Abstract: Recent vision language models (VLMs) have made remarkable strides in generative modeling with multimodal inputs, particularly text and images. Howev

26 May 2026

Approximating Safety Feedback Without a Safety Oracle via Model Predictive Control

SafetyDGX agent

arXiv:2510.20955v2 Announce Type: replace Abstract: Safe decision-making algorithms for control of mobile robots often require the existence of feedback to verify the safety of proposed actions. This

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

SafetyDGX agent

arXiv:2505.18979v2 Announce Type: replace Abstract: Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters.

Safety-Critical Whole-Body Control for Humanoid Robots via Input-to-State Safe Control Barrier Functions

SafetyDGX agent

arXiv:2605.25546v1 Announce Type: new Abstract: Safety-critical control is essential for humanoid robots operating in complex human-centered environments, where physical safety constraints such as joi

18 May 2026

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

Model ReleasesDGX agent

arXiv:2605.15239v1 Announce Type: new Abstract: Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is di

14 May 2026

Distributionally Robust Safety Under Arbitrary Uncertainties: A Safety Filtering Approach

SafetyDGX agent

arXiv:2605.12974v1 Announce Type: new Abstract: In this work, we study how to ensure probabilistic safety for nonlinear systems under distributional ambiguity. Our approach builds on a backup-based sa

Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements

SafetyDGX agent

arXiv:2605.12963v1 Announce Type: new Abstract: As AI systems become increasingly capable, safety strategies must be evaluated not only by how much they reduce present risk, but by whether they could

21 Apr 2026

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? …

SafetyDGX agent

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? This link is from Nov 2023 & shows how tech CEOs used 'safet

Safe Control using Learned Safety Filters and Adaptive Conformal Inference

Model ReleasesDGX agent

arXiv:2604.18482v1 Announce Type: cross Abstract: Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability chal

11 Jun 2026

AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin

Model ReleasesDGX agent

arXiv:2506.08473v4 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely

25 May 2026

A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies (New York Times)

SafetyDGX agent

New York Times: A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies — The government's A.I. Se

13 May 2026

OpenAI endorses the Kids Online Safety Act and Illinois SB 315, an AI safety bill to create requirements around transparency, incident reporting, and more (OpenAI Global Affairs)

SafetyDGX agent

OpenAI Global Affairs: OpenAI endorses the Kids Online Safety Act and Illinois SB 315, an AI safety bill to create requirements around transparency, incident reporting, and more — Welcome (back) to Th

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing

SafetyDGX agent

arXiv:2602.02280v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face severe safety risks from jailbreak attacks, yet current safety testing largely relies on static datasets and

16 Apr 2026

Boundary Sampling to Learn Predictive Safety Filters via Pontryagin's Maximum Principle

SafetyDGX agent

arXiv:2604.13325v1 Announce Type: new Abstract: Safety filters provide a practical approach for enforcing safety constraints in autonomous systems. While learning-based tools scale to high-dimensional

9 Jul 2026

Initiation Safety: A Missing Dimension in Generalist-Robot Safety

SafetyDGX agent

arXiv:2607.07420v1 Announce Type: new Abstract: Safety for generalist robots is usually discussed in terms of motion or dialogue. We argue a third question is missing: should the robot take its first

27 Jul 2026

Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

SafetyDGX agent

arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmen

28 May 2026

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Model ReleasesDGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

4 May 2026

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

SafetyDGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

14 Apr 2026

Dialogue based Interactive Explanations for Safety Decisions in Human Robot Collaboration

SafetyDGX agent

arXiv:2604.05896v2 Announce Type: replace Abstract: As robots increasingly operate in shared, safety critical environments, acting safely is no longer sufficient robots must also make their safety dec

12 Aug 2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

SafetyDGX agent

arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formula

Dual Stress: Runtime Safety Monitoring for Safety-Constrained MPC Navigation

SafetyDGX agent

arXiv:2608.10791v1 Announce Type: new Abstract: Runtime hazard monitors for autonomous naviga- tion are conventionally built from geometric quantities: predicted clearance, time to collision, and requ

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

SafetyDGX agent

arXiv:2608.10513v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their

23 Jul 2026

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

SafetyDGX agent

arXiv:2607.19701v1 Announce Type: new Abstract: VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among the

23 Jun 2026

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.23686v1 Announce Type: new Abstract: Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains large

9 Jun 2026

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

Model ReleasesDGX agent

arXiv:2606.09778v1 Announce Type: cross Abstract: Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered control

5 Jun 2026

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

SafetyDGX agent

arXiv:2606.05290v1 Announce Type: new Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring ret

20 May 2026

Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints

SafetyDGX agent

arXiv:2605.18842v1 Announce Type: new Abstract: Safe reinforcement learning in nonstationary environments requires safety mechanisms that adapt as environmental conditions change. Standard safe reinfo

From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning

SafetyDGX agent

arXiv:2605.18841v1 Announce Type: new Abstract: Safety in reinforcement learning is often specified through cumulative cost constraints, but these trajectory-level guarantees do not directly prevent u

30 Apr 2026

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

Model ReleasesDGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

4 Aug 2026

Mistral releases Shieldstral, a 3B multimodal safety classifier that it says matches models up to 7x its size on text safety, available under Apache 2.0 (Mistral AI Blog)

Model ReleasesDGX agent

Mistral AI Blog: Mistral releases Shieldstral, a 3B multimodal safety classifier that it says matches models up to 7x its size on text safety, available under Apache 2.0 — Every product that ships a m

2 Aug 2026

A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety (Ben Cohen/Wall Street Journal)

SafetyDGX agent

Ben Cohen / Wall Street Journal: A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety — Jacob Tsim

1 Jun 2026

Configurable Reward Model for Balanced Safety Alignment

SafetyDGX agent

arXiv:2605.30487v1 Announce Type: new Abstract: Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned

19 May 2026

Why Do Safety Guardrails Degrade Across Languages?

SafetyDGX agent

arXiv:2605.17173v1 Announce Type: cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds

Policy Library CBF: Finite-Horizon Safety at Runtime via Parallel Rollouts

SafetyDGX agent

arXiv:2605.16588v1 Announce Type: new Abstract: Safety-critical autonomy in unstructured environments poses significant challenges for online safety certification under evolving constraints. We propos

12 May 2026

The Safety-Aware Denoiser for Text Diffusion Models

SafetyDGX agent

arXiv:2605.08116v1 Announce Type: cross Abstract: Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored.

29 Apr 2026

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty

SafetyDGX agent

arXiv:2604.25508v1 Announce Type: new Abstract: Safety remains an open problem in reinforcement learning (RL), especially during training. While safety filters are promising to address safe exploratio

31 Jul 2026

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

SafetyDGX agent

arXiv:2603.15684v2 Announce Type: replace-cross Abstract: Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although

15 Jul 2026

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

SafetyDGX agent

arXiv:2510.13698v4 Announce Type: replace Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety

8 Jul 2026

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

SafetyDGX agent

arXiv:2607.06326v1 Announce Type: new Abstract: Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-l

26 Jun 2026

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.27079v1 Announce Type: new Abstract: In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models cont

6 Jun 2026

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

Model ReleasesDGX agent

arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and r

7 May 2026

Worth remembering when @janleike quit over OpenAI safety concerns.

SafetyDGX agent

Jan Leike, OpenAI's safety and alignment lead, resigned in May 2024, citing concerns about the company's commitment to safety practices and the deprioritization of safety work relative to product deve

6 May 2026

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models

SafetyDGX agent

arXiv:2605.02914v1 Announce Type: new Abstract: A guard model fine-tuned on entirely benign data can lose all safety alignment -- not through adversarial manipulation, but through standard domain spec

5 May 2026

To Do or Not to Do: Ensuring the Safety of Visuomotor Policies Learned from Demonstrations

SafetyDGX agent

arXiv:2605.01201v1 Announce Type: new Abstract: Task success has historically been the primary measure of policy performance in imitation learning (IL) research. This characteristics strictly limits t

← Previous
123…238
Next →
14,230 results