AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
49+ results
Safety

Full-Body Dynamic Safety for Robot Manipulators: 3D Poisson Safety Functions for CBF-Based Safety Filters

DGX agent

arXiv:2604.21189v1 Announce Type: new Abstract: Collision avoidance for robotic manipulators requires enforcing full-body safety constraints in high-dimensional configuration spaces. Control Barrier F

safetyarxiv-cs-ro
24 Apr 2026
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Deep QP Safety Filter: Model-free Learning for Reachability-based Safety Filter

DGX agent

arXiv:2601.21297v2 Announce Type: replace Abstract: We introduce Deep QP Safety Filter, a fully data-driven safety layer for black-box dynamical systems. Our method learns a Quadratic-Program (QP) saf

safetyarxiv-cs-ro
15 Apr 2026
Model Releases

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

DGX agent

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, I

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

DGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

safetyarxiv-cs-cl
4 Jun 2026
Safety

Cooptimizing Safety and Performance Using Safety Value-Constrained Model Predictive Control

DGX agent

arXiv:2604.23863v1 Announce Type: new Abstract: Autonomous systems are increasingly deployed in real-world environments, where they must achieve high performance while maintaining safety under state a

safetyarxiv-cs-ro
28 Apr 2026
Local Ai

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

DGX agent

arXiv:2608.09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial i

local-aiarxiv-cs-ai
11 Aug 2026
Safety

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

DGX agent

arXiv:2604.00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Mod

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

Permissive Safety Through Trusted Inference: Verifiable Belief-Space Neural Safety Filters for Assured Interactive Robotics

DGX agent

arXiv:2606.02562v1 Announce Type: cross Abstract: Autonomous robots that interact with people must make safe and efficient decisions under human-induced uncertainty, such as their preferences, goals,

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Approximating Safety Feedback Without a Safety Oracle via Model Predictive Control

DGX agent

arXiv:2510.20955v2 Announce Type: replace Abstract: Safe decision-making algorithms for control of mobile robots often require the existence of feedback to verify the safety of proposed actions. This

safetyarxiv-cs-lg
26 May 2026
Model Releases

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

DGX agent

arXiv:2605.15239v1 Announce Type: new Abstract: Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is di

model-releasesarxiv-cs-lg
18 May 2026
Safety

Distributionally Robust Safety Under Arbitrary Uncertainties: A Safety Filtering Approach

DGX agent

arXiv:2605.12974v1 Announce Type: new Abstract: In this work, we study how to ensure probabilistic safety for nonlinear systems under distributional ambiguity. Our approach builds on a backup-based sa

safetyarxiv-cs-ro
14 May 2026
Model Releases

AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin

DGX agent

arXiv:2506.08473v4 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely

model-releasesarxiv-cs-lg
11 Jun 2026
Safety

Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

DGX agent

arXiv:2503.11832v5 Announce Type: replace Abstract: Recent vision language models (VLMs) have made remarkable strides in generative modeling with multimodal inputs, particularly text and images. Howev

safetyarxiv-cs-ai
2 Jun 2026
Safety

Safety-aware Goal-oriented Semantic Sensing, Communication, and Control for Robotics

DGX agent

arXiv:2603.13502v2 Announce Type: replace Abstract: Wirelessly-connected robotic systems empower robots with real-time intelligence by leveraging remote computing resources for decision-making. Howeve

safetyarxiv-cs-ro
28 Apr 2026
Safety

Boundary Sampling to Learn Predictive Safety Filters via Pontryagin's Maximum Principle

DGX agent

arXiv:2604.13325v1 Announce Type: new Abstract: Safety filters provide a practical approach for enforcing safety constraints in autonomous systems. While learning-based tools scale to high-dimensional

safetyarxiv-cs-ro
16 Apr 2026
Safety

Initiation Safety: A Missing Dimension in Generalist-Robot Safety

DGX agent

arXiv:2607.07420v1 Announce Type: new Abstract: Safety for generalist robots is usually discussed in terms of motion or dialogue. We argue a third question is missing: should the robot take its first

safetyarxiv-cs-ro
9 Jul 2026
Safety

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

DGX agent

arXiv:2505.18979v2 Announce Type: replace Abstract: Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters.

safetyarxiv-cs-lg
26 May 2026
Safety

Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

DGX agent

arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmen

safetyarxiv-cs-lg
27 Jul 2026
Model Releases

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

DGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

model-releasesarxiv-cs-ai
28 May 2026
Safety

Safety-Critical Whole-Body Control for Humanoid Robots via Input-to-State Safe Control Barrier Functions

DGX agent

arXiv:2605.25546v1 Announce Type: new Abstract: Safety-critical control is essential for humanoid robots operating in complex human-centered environments, where physical safety constraints such as joi

safetyarxiv-cs-ro
26 May 2026
Safety

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

DGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

safetyarxiv-cs-cl
4 May 2026
Model Releases

Safe Control using Learned Safety Filters and Adaptive Conformal Inference

DGX agent

arXiv:2604.18482v1 Announce Type: cross Abstract: Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability chal

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

Dialogue based Interactive Explanations for Safety Decisions in Human Robot Collaboration

DGX agent

arXiv:2604.05896v2 Announce Type: replace Abstract: As robots increasingly operate in shared, safety critical environments, acting safely is no longer sufficient robots must also make their safety dec

safetyarxiv-cs-ro
14 Apr 2026
Safety

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

DGX agent

arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formula

safetyarxiv-cs-lg
12 Aug 2026
Safety

Beyond 'I Can't Help With That': How Child Safety Experts Evaluate AI Chatbot Safety

DGX agent

arXiv:2608.07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes s

safetyarxiv-cs-ai
11 Aug 2026
Safety

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

DGX agent

arXiv:2607.19701v1 Announce Type: new Abstract: VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among the

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

DGX agent

arXiv:2606.23686v1 Announce Type: new Abstract: Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains large

model-releasesarxiv-cs-ro
23 Jun 2026
Model Releases

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

DGX agent

arXiv:2606.09778v1 Announce Type: cross Abstract: Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered control

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

DGX agent

arXiv:2606.05290v1 Announce Type: new Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring ret

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

DGX agent

arXiv:2603.10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints

DGX agent

arXiv:2605.18842v1 Announce Type: new Abstract: Safe reinforcement learning in nonstationary environments requires safety mechanisms that adapt as environmental conditions change. Standard safe reinfo

safetyarxiv-cs-lg
20 May 2026
Model Releases

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

DGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

model-releasesarxiv-cs-ai
30 Apr 2026
Safety

Dual Stress: Runtime Safety Monitoring for Safety-Constrained MPC Navigation

DGX agent

arXiv:2608.10791v1 Announce Type: new Abstract: Runtime hazard monitors for autonomous naviga- tion are conventionally built from geometric quantities: predicted clearance, time to collision, and requ

safetyarxiv-cs-ro
12 Aug 2026
Safety

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

DGX agent

arXiv:2608.10513v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their

safetyarxiv-cs-ai
12 Aug 2026
Safety

Configurable Reward Model for Balanced Safety Alignment

DGX agent

arXiv:2605.30487v1 Announce Type: new Abstract: Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned

safetyarxiv-cs-cl
1 Jun 2026
Safety

Why Do Safety Guardrails Degrade Across Languages?

DGX agent

arXiv:2605.17173v1 Announce Type: cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds

safetyarxiv-cs-ai
19 May 2026
Safety

Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements

DGX agent

arXiv:2605.12963v1 Announce Type: new Abstract: As AI systems become increasingly capable, safety strategies must be evaluated not only by how much they reduce present risk, but by whether they could

safetyarxiv-cs-ai
14 May 2026
Safety

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing

DGX agent

arXiv:2602.02280v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face severe safety risks from jailbreak attacks, yet current safety testing largely relies on static datasets and

safetyarxiv-cs-cl
13 May 2026
Safety

The Safety-Aware Denoiser for Text Diffusion Models

DGX agent

arXiv:2605.08116v1 Announce Type: cross Abstract: Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored.

safetyarxiv-cs-ai
12 May 2026
Safety

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty

DGX agent

arXiv:2604.25508v1 Announce Type: new Abstract: Safety remains an open problem in reinforcement learning (RL), especially during training. While safety filters are promising to address safe exploratio

safetyarxiv-cs-lg
29 Apr 2026
Safety

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

DGX agent

arXiv:2603.15684v2 Announce Type: replace-cross Abstract: Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although

safetyarxiv-cs-ai
31 Jul 2026
Safety

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

DGX agent

arXiv:2510.13698v4 Announce Type: replace Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety

safetyarxiv-cs-cv
15 Jul 2026
Safety

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

DGX agent

arXiv:2607.06326v1 Announce Type: new Abstract: Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-l

safetyarxiv-cs-ai
8 Jul 2026
Model Releases

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

DGX agent

arXiv:2606.27079v1 Announce Type: new Abstract: In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models cont

model-releasesarxiv-cs-ro
26 Jun 2026
Model Releases

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

DGX agent

arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and r

model-releasesarxiv-cs-ai
6 Jun 2026
Safety

COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection

DGX agent

arXiv:2606.04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-valu

safetyarxiv-cs-lg
4 Jun 2026
Safety

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

DGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

safetyarxiv-cs-ai
4 Jun 2026
Safety

From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning

DGX agent

arXiv:2605.18841v1 Announce Type: new Abstract: Safety in reinforcement learning is often specified through cumulative cost constraints, but these trajectory-level guarantees do not directly prevent u

safetyarxiv-cs-lg
20 May 2026
← Previous
1
Next →
12,202 results
← Previous
123…255
Next →