AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlog
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,230 results
Safety

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

DGX agent

arXiv:2510.13698v4 Announce Type: replace Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety

safetyarxiv-cs-cv
15 Jul 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

DGX agent

arXiv:2607.06326v1 Announce Type: new Abstract: Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-l

safetyarxiv-cs-ai
8 Jul 2026
Model Releases

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

DGX agent

arXiv:2606.27079v1 Announce Type: new Abstract: In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models cont

model-releasesarxiv-cs-ro
26 Jun 2026
Model Releases

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

DGX agent

arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and r

model-releasesarxiv-cs-ai
6 Jun 2026
Safety

COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection

DGX agent

arXiv:2606.04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-valu

safetyarxiv-cs-lg
4 Jun 2026
Safety

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

DGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

safetyarxiv-cs-ai
4 Jun 2026
Safety

From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning

DGX agent

arXiv:2605.18841v1 Announce Type: new Abstract: Safety in reinforcement learning is often specified through cumulative cost constraints, but these trajectory-level guarantees do not directly prevent u

safetyarxiv-cs-lg
20 May 2026
Safety

Policy Library CBF: Finite-Horizon Safety at Runtime via Parallel Rollouts

DGX agent

arXiv:2605.16588v1 Announce Type: new Abstract: Safety-critical autonomy in unstructured environments poses significant challenges for online safety certification under evolving constraints. We propos

safetyarxiv-cs-ro
19 May 2026
Safety

Worth remembering when @janleike quit over OpenAI safety concerns.

DGX agent

Jan Leike, OpenAI's safety and alignment lead, resigned in May 2024, citing concerns about the company's commitment to safety practices and the deprioritization of safety work relative to product deve

safetygary-marcus--x
7 May 2026
Safety

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models

DGX agent

arXiv:2605.02914v1 Announce Type: new Abstract: A guard model fine-tuned on entirely benign data can lose all safety alignment -- not through adversarial manipulation, but through standard domain spec

safetyarxiv-cs-lg
6 May 2026
Safety

To Do or Not to Do: Ensuring the Safety of Visuomotor Policies Learned from Demonstrations

DGX agent

arXiv:2605.01201v1 Announce Type: new Abstract: Task success has historically been the primary measure of policy performance in imitation learning (IL) research. This characteristics strictly limits t

safetyarxiv-cs-ro
5 May 2026
Safety

Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

DGX agent

arXiv:2604.23775v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of safety challenges, s

safetyarxiv-cs-ro
28 Apr 2026
Safety

Dataset Safety in Autonomous Driving: Requirements, Risks, and Assurance

DGX agent

arXiv:2511.08439v2 Announce Type: replace Abstract: Dataset integrity is fundamental to the safety and reliability of AI systems, especially in autonomous driving. This paper presents a structured fra

safetyarxiv-cs-ai
15 Apr 2026
Safety

Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model

DGX agent

arXiv:2604.09665v1 Announce Type: cross Abstract: While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted

safetyarxiv-cs-ai
14 Apr 2026
Safety

Online Learning-Enhanced High Order Adaptive Safety Control

DGX agent

arXiv:2511.19651v2 Announce Type: replace Abstract: Control barrier functions (CBFs) are an effective model-based tool to formally certify the safety of a system. With the growing complexity of modern

safetyarxiv-cs-ro
14 Apr 2026
Safety

Towards provable probabilistic safety for scalable embodied AI systems

DGX agent

arXiv:2506.05171v3 Announce Type: replace-cross Abstract: Embodied AI systems, comprising AI models and physical plants, are increasingly prevalent across various applications. Due to the rarity of sy

safetyarxiv-cs-ai
10 Apr 2026
Safety

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

DGX agent

arXiv:2608.02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert

safetyarxiv-cs-ai
5 Aug 2026
Safety

Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models

DGX agent

arXiv:2511.21214v4 Announce Type: replace-cross Abstract: Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We stu

safetyarxiv-cs-ai
5 Aug 2026
Local Ai

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

DGX agent

arXiv:2607.28969v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes

local-aiarxiv-cs-cv
3 Aug 2026
Model Releases

SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems

DGX agent

arXiv:2603.03536v2 Announce Type: replace-cross Abstract: Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an un

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

DGX agent

arXiv:2607.02914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness,

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

DGX agent

arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testin

model-releasesarxiv-cs-ai
3 Jul 2026
Safety

Safe Online Learning via Smooth Safety-Structured Policy Composition

DGX agent

arXiv:2606.31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics. Existing approaches t

safetyarxiv-cs-lg
1 Jul 2026
Safety

Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

DGX agent

arXiv:2601.12033v2 Announce Type: replace Abstract: Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety,

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

DGX agent

arXiv:2606.27632v1 Announce Type: new Abstract: As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We arg

model-releasesarxiv-cs-cl
29 Jun 2026
Safety

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

DGX agent

arXiv:2606.08234v1 Announce Type: new Abstract: LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning

safetyarxiv-cs-ai
9 Jun 2026
Safety

Extending Responsibility-Sensitive Safety for the Assessment of Offloaded Autonomous Driving Services

DGX agent

arXiv:2606.07067v1 Announce Type: new Abstract: Safety is a fundamental requirement in the development of autonomous driving (AD) systems. While function offloading has demonstrated significant benefi

safetyarxiv-cs-ro
8 Jun 2026
Safety

SafeGene: Reusable Adapters for Transferable Safety Alignment

DGX agent

arXiv:2606.06519v1 Announce Type: new Abstract: Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vul

safetyarxiv-cs-ai
8 Jun 2026
Safety

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

DGX agent

arXiv:2606.05743v1 Announce Type: cross Abstract: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifi

safetyarxiv-cs-cl
5 Jun 2026
Safety

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

DGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

safetyarxiv-cs-ai
29 May 2026
Safety

Synergistic Simplex: Cooperative Runtime Assurance for Safety-Critical Autonomous Systems

DGX agent

arXiv:2605.08190v1 Announce Type: new Abstract: Autonomous systems increasingly rely on machine-learning (ML) components for safety-critical tasks such as perception and control in autonomous vehicles

safetyarxiv-cs-lg
12 May 2026
Safety

Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation

DGX agent

arXiv:2603.17368v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning cap

safetyarxiv-cs-ai
6 May 2026
Safety

Policy-Grounded Safety Evaluation of 20 Large Language Models

DGX agent

arXiv:2507.14719v2 Announce Type: replace Abstract: As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. T

safetyarxiv-cs-ai
1 May 2026
Safety

Discovering Agentic Safety Specifications from 1-Bit Danger Signals

DGX agent

arXiv:2604.23210v1 Announce Type: new Abstract: Can large language model agents discover hidden safety objectives through experience alone? We introduce EPO-Safe (Experiential Prompt Optimization for

safetyarxiv-cs-ai
28 Apr 2026
Safety

When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models

DGX agent

arXiv:2510.21285v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex multi-step reasoning, yet they still exhibit severe safety failures such as harm

safetyarxiv-cs-ai
27 Apr 2026
Model Releases

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

DGX agent

arXiv:2503.03480v4 Announce Type: replace Abstract: Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-w

model-releasesarxiv-cs-ro
21 Apr 2026
Safety

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

DGX agent

arXiv:2603.13933v2 Announce Type: replace Abstract: Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on a

safetyarxiv-cs-cl
17 Apr 2026
Safety

RL-STPA: Adapting System-Theoretic Hazard Analysis for Safety-Critical Reinforcement Learning

DGX agent

arXiv:2604.15201v1 Announce Type: new Abstract: As reinforcement learning (RL) deployments expand into safety-critical domains, existing evaluation methods fail to systematically identify hazards aris

safetyarxiv-cs-lg
17 Apr 2026
Safety

Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems

DGX agent

arXiv:2608.02806v1 Announce Type: cross Abstract: By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve roa

safetyarxiv-cs-cv
5 Aug 2026
Safety

Shielding for Higher-Order Safety

DGX agent

arXiv:2608.03662v1 Announce Type: new Abstract: Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety. Classical shields are usually synthesis

safetyarxiv-cs-ai
5 Aug 2026
Safety

Epistemic Norms for AI Safety and Alignment Research

DGX agent

arXiv:2607.24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment resea

safetyarxiv-cs-ai
28 Jul 2026
Safety

Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards

DGX agent

arXiv:2607.19827v1 Announce Type: new Abstract: Ensuring safety in Physical AI systems operating in real-world environments is a critical challenge, particularly in hospital wards where vulnerable pat

safetyarxiv-cs-ro
23 Jul 2026
Safety

Designing Safety-Constrained LLM Systems for Public Health Information Access

DGX agent

arXiv:2607.13038v1 Announce Type: cross Abstract: We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on m

safetyarxiv-cs-ai
16 Jul 2026
Safety

Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning

DGX agent

arXiv:2602.13562v2 Announce Type: replace-cross Abstract: While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measu

safetyarxiv-cs-ai
30 Jun 2026
Safety

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

DGX agent

arXiv:2606.07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement

safetyarxiv-cs-ai
9 Jun 2026
Safety

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

DGX agent

arXiv:2606.06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework f

safetyarxiv-cs-ai
8 Jun 2026
Safety

When Autoregressive Consistency Hurts Safety Alignment

DGX agent

arXiv:2606.04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near t

safetyarxiv-cs-lg
4 Jun 2026
Safety

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

DGX agent

arXiv:2606.03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identifica

safetyarxiv-cs-ai
3 Jun 2026
← Previous
1234…297
Next →