AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
49+ results
Safety

Full-Body Dynamic Safety for Robot Manipulators: 3D Poisson Safety Functions for CBF-Based Safety Filters

DGX agent

arXiv:2604.21189v1 Announce Type: new Abstract: Collision avoidance for robotic manipulators requires enforcing full-body safety constraints in high-dimensional configuration spaces. Control Barrier F

safetyarxiv-cs-ro
24 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

OpenAI's head of safety, Johannes Heidecke, is leaving as OpenAI integrates its research and safety teams; Mia Glaese will become VP of research and safety (Maxwell Zeff/Wired)

DGX agent

Maxwell Zeff / Wired: OpenAI's head of safety, Johannes Heidecke, is leaving as OpenAI integrates its research and safety teams; Mia Glaese will become VP of research and safety — Johannes Heidecke's

safetytechmeme
10 Jul 2026
Safety

Deep QP Safety Filter: Model-free Learning for Reachability-based Safety Filter

DGX agent

arXiv:2601.21297v2 Announce Type: replace Abstract: We introduce Deep QP Safety Filter, a fully data-driven safety layer for black-box dynamical systems. Our method learns a Quadratic-Program (QP) saf

safetyarxiv-cs-ro
15 Apr 2026
Model Releases

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

DGX agent

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, I

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

DGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

safetyarxiv-cs-cl
4 Jun 2026
Safety

Cooptimizing Safety and Performance Using Safety Value-Constrained Model Predictive Control

DGX agent

arXiv:2604.23863v1 Announce Type: new Abstract: Autonomous systems are increasingly deployed in real-world environments, where they must achieve high performance while maintaining safety under state a

safetyarxiv-cs-ro
28 Apr 2026
Local Ai

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

DGX agent

arXiv:2608.09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial i

local-aiarxiv-cs-ai
11 Aug 2026
Safety

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

DGX agent

arXiv:2604.00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Mod

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

Permissive Safety Through Trusted Inference: Verifiable Belief-Space Neural Safety Filters for Assured Interactive Robotics

DGX agent

arXiv:2606.02562v1 Announce Type: cross Abstract: Autonomous robots that interact with people must make safe and efficient decisions under human-induced uncertainty, such as their preferences, goals,

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Approximating Safety Feedback Without a Safety Oracle via Model Predictive Control

DGX agent

arXiv:2510.20955v2 Announce Type: replace Abstract: Safe decision-making algorithms for control of mobile robots often require the existence of feedback to verify the safety of proposed actions. This

safetyarxiv-cs-lg
26 May 2026
Model Releases

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

DGX agent

arXiv:2605.15239v1 Announce Type: new Abstract: Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is di

model-releasesarxiv-cs-lg
18 May 2026
Safety

Distributionally Robust Safety Under Arbitrary Uncertainties: A Safety Filtering Approach

DGX agent

arXiv:2605.12974v1 Announce Type: new Abstract: In this work, we study how to ensure probabilistic safety for nonlinear systems under distributional ambiguity. Our approach builds on a backup-based sa

safetyarxiv-cs-ro
14 May 2026
Safety

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? …

DGX agent

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? This link is from Nov 2023 & shows how tech CEOs used 'safet

safetygary-marcus--x
21 Apr 2026
Model Releases

AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin

DGX agent

arXiv:2506.08473v4 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely

model-releasesarxiv-cs-lg
11 Jun 2026
Safety

Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

DGX agent

arXiv:2503.11832v5 Announce Type: replace Abstract: Recent vision language models (VLMs) have made remarkable strides in generative modeling with multimodal inputs, particularly text and images. Howev

safetyarxiv-cs-ai
2 Jun 2026
Safety

A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies (New York Times)

DGX agent

New York Times: A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies — The government's A.I. Se

safetytechmeme
25 May 2026
Safety

OpenAI endorses the Kids Online Safety Act and Illinois SB 315, an AI safety bill to create requirements around transparency, incident reporting, and more (OpenAI Global Affairs)

DGX agent

OpenAI Global Affairs: OpenAI endorses the Kids Online Safety Act and Illinois SB 315, an AI safety bill to create requirements around transparency, incident reporting, and more — Welcome (back) to Th

safetytechmeme
13 May 2026
Safety

Safety-aware Goal-oriented Semantic Sensing, Communication, and Control for Robotics

DGX agent

arXiv:2603.13502v2 Announce Type: replace Abstract: Wirelessly-connected robotic systems empower robots with real-time intelligence by leveraging remote computing resources for decision-making. Howeve

safetyarxiv-cs-ro
28 Apr 2026
Safety

Boundary Sampling to Learn Predictive Safety Filters via Pontryagin's Maximum Principle

DGX agent

arXiv:2604.13325v1 Announce Type: new Abstract: Safety filters provide a practical approach for enforcing safety constraints in autonomous systems. While learning-based tools scale to high-dimensional

safetyarxiv-cs-ro
16 Apr 2026
Safety

Initiation Safety: A Missing Dimension in Generalist-Robot Safety

DGX agent

arXiv:2607.07420v1 Announce Type: new Abstract: Safety for generalist robots is usually discussed in terms of motion or dialogue. We argue a third question is missing: should the robot take its first

safetyarxiv-cs-ro
9 Jul 2026
Safety

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

DGX agent

arXiv:2505.18979v2 Announce Type: replace Abstract: Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters.

safetyarxiv-cs-lg
26 May 2026
Safety

Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

DGX agent

arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmen

safetyarxiv-cs-lg
27 Jul 2026
Model Releases

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

DGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

model-releasesarxiv-cs-ai
28 May 2026
Safety

Safety-Critical Whole-Body Control for Humanoid Robots via Input-to-State Safe Control Barrier Functions

DGX agent

arXiv:2605.25546v1 Announce Type: new Abstract: Safety-critical control is essential for humanoid robots operating in complex human-centered environments, where physical safety constraints such as joi

safetyarxiv-cs-ro
26 May 2026
Safety

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

DGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

safetyarxiv-cs-cl
4 May 2026
Model Releases

Safe Control using Learned Safety Filters and Adaptive Conformal Inference

DGX agent

arXiv:2604.18482v1 Announce Type: cross Abstract: Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability chal

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

Dialogue based Interactive Explanations for Safety Decisions in Human Robot Collaboration

DGX agent

arXiv:2604.05896v2 Announce Type: replace Abstract: As robots increasingly operate in shared, safety critical environments, acting safely is no longer sufficient robots must also make their safety dec

safetyarxiv-cs-ro
14 Apr 2026
Safety

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

DGX agent

arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formula

safetyarxiv-cs-lg
12 Aug 2026
Safety

Beyond 'I Can't Help With That': How Child Safety Experts Evaluate AI Chatbot Safety

DGX agent

arXiv:2608.07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes s

safetyarxiv-cs-ai
11 Aug 2026
Safety

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

DGX agent

arXiv:2607.19701v1 Announce Type: new Abstract: VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among the

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

DGX agent

arXiv:2606.23686v1 Announce Type: new Abstract: Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains large

model-releasesarxiv-cs-ro
23 Jun 2026
Model Releases

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

DGX agent

arXiv:2606.09778v1 Announce Type: cross Abstract: Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered control

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

DGX agent

arXiv:2606.05290v1 Announce Type: new Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring ret

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

DGX agent

Nemotron 3.5 Content Safety is NVIDIA's multimodal safety solution designed for enterprise AI applications, offering customizable safeguards for both text and image inputs across different global cont

model-releaseshugging-face
4 Jun 2026
Model Releases

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

DGX agent

arXiv:2603.10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints

DGX agent

arXiv:2605.18842v1 Announce Type: new Abstract: Safe reinforcement learning in nonstationary environments requires safety mechanisms that adapt as environmental conditions change. Standard safe reinfo

safetyarxiv-cs-lg
20 May 2026
Model Releases

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

DGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

model-releasesarxiv-cs-ai
30 Apr 2026
Safety

Dual Stress: Runtime Safety Monitoring for Safety-Constrained MPC Navigation

DGX agent

arXiv:2608.10791v1 Announce Type: new Abstract: Runtime hazard monitors for autonomous naviga- tion are conventionally built from geometric quantities: predicted clearance, time to collision, and requ

safetyarxiv-cs-ro
12 Aug 2026
Safety

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

DGX agent

arXiv:2608.10513v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

Mistral releases Shieldstral, a 3B multimodal safety classifier that it says matches models up to 7x its size on text safety, available under Apache 2.0 (Mistral AI Blog)

DGX agent

Mistral AI Blog: Mistral releases Shieldstral, a 3B multimodal safety classifier that it says matches models up to 7x its size on text safety, available under Apache 2.0 — Every product that ships a m

model-releasestechmeme
4 Aug 2026
Safety

A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety (Ben Cohen/Wall Street Journal)

DGX agent

Ben Cohen / Wall Street Journal: A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety — Jacob Tsim

safetytechmeme
2 Aug 2026
Safety

Configurable Reward Model for Balanced Safety Alignment

DGX agent

arXiv:2605.30487v1 Announce Type: new Abstract: Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned

safetyarxiv-cs-cl
1 Jun 2026
Safety

Why Do Safety Guardrails Degrade Across Languages?

DGX agent

arXiv:2605.17173v1 Announce Type: cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds

safetyarxiv-cs-ai
19 May 2026
Safety

Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements

DGX agent

arXiv:2605.12963v1 Announce Type: new Abstract: As AI systems become increasingly capable, safety strategies must be evaluated not only by how much they reduce present risk, but by whether they could

safetyarxiv-cs-ai
14 May 2026
Safety

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing

DGX agent

arXiv:2602.02280v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face severe safety risks from jailbreak attacks, yet current safety testing largely relies on static datasets and

safetyarxiv-cs-cl
13 May 2026
Safety

The Safety-Aware Denoiser for Text Diffusion Models

DGX agent

arXiv:2605.08116v1 Announce Type: cross Abstract: Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored.

safetyarxiv-cs-ai
12 May 2026
Safety

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty

DGX agent

arXiv:2604.25508v1 Announce Type: new Abstract: Safety remains an open problem in reinforcement learning (RL), especially during training. While safety filters are promising to address safe exploratio

safetyarxiv-cs-lg
29 Apr 2026
Safety

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

DGX agent

arXiv:2603.15684v2 Announce Type: replace-cross Abstract: Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although

safetyarxiv-cs-ai
31 Jul 2026
← Previous
123…297
Next →
14,234 results