AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,230 results
15 Apr 2026

Dataset Safety in Autonomous Driving: Requirements, Risks, and Assurance

SafetyDGX agent

arXiv:2511.08439v2 Announce Type: replace Abstract: Dataset integrity is fundamental to the safety and reliability of AI systems, especially in autonomous driving. This paper presents a structured fra

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints

SafetyDGX agent

arXiv:2604.12384v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) remains highly fragile during fine-tuning, where even benign adaptation can degrade pre-trained refusal

14 Apr 2026

Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

arXiv:2604.09665v1 Announce Type: cross Abstract: While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted

Online Learning-Enhanced High Order Adaptive Safety Control

SafetyDGX agent

arXiv:2511.19651v2 Announce Type: replace Abstract: Control barrier functions (CBFs) are an effective model-based tool to formally certify the safety of a system. With the growing complexity of modern

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score…

SafetyDGX agent

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score of 100. This allows you to maintain a higher average safety

10 Apr 2026

Towards provable probabilistic safety for scalable embodied AI systems

SafetyDGX agent

arXiv:2506.05171v3 Announce Type: replace-cross Abstract: Embodied AI systems, comprising AI models and physical plants, are increasingly prevalent across various applications. Due to the rarity of sy

5 Aug 2026

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

SafetyDGX agent

arXiv:2608.02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert

Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models

SafetyDGX agent

arXiv:2511.21214v4 Announce Type: replace-cross Abstract: Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We stu

Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems

SafetyDGX agent

arXiv:2608.02806v1 Announce Type: cross Abstract: By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve roa

Shielding for Higher-Order Safety

SafetyDGX agent

arXiv:2608.03662v1 Announce Type: new Abstract: Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety. Classical shields are usually synthesis

White House, AI firms keep safety framework talks private

SafetyDGX agent

The White House met with representatives from leading artificial intelligence companies today to discuss a safety framework for the government to review frontier models prior to launch, although there

3 Aug 2026

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

Local AiDGX agent

arXiv:2607.28969v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes

28 Jul 2026

SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems

Model ReleasesDGX agent

arXiv:2603.03536v2 Announce Type: replace-cross Abstract: Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an un

Epistemic Norms for AI Safety and Alignment Research

SafetyDGX agent

arXiv:2607.24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment resea

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

Model ReleasesDGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

7 Jul 2026

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

SafetyDGX agent

arXiv:2607.02914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness,

3 Jul 2026

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Model ReleasesDGX agent

arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testin

1 Jul 2026

Safe Online Learning via Smooth Safety-Structured Policy Composition

SafetyDGX agent

arXiv:2606.31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics. Existing approaches t

Harnessing Textual Refusal Directions for Multimodal Safety

SafetyDGX agent

arXiv:2606.31876v1 Announce Type: new Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. B

30 Jun 2026

Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

SafetyDGX agent

arXiv:2601.12033v2 Announce Type: replace Abstract: Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety,

Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning

SafetyDGX agent

arXiv:2602.13562v2 Announce Type: replace-cross Abstract: While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measu

29 Jun 2026

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

Model ReleasesDGX agent

arXiv:2606.27632v1 Announce Type: new Abstract: As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We arg

9 Jun 2026

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

SafetyDGX agent

arXiv:2606.08234v1 Announce Type: new Abstract: LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

SafetyDGX agent

arXiv:2606.07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement

8 Jun 2026

Extending Responsibility-Sensitive Safety for the Assessment of Offloaded Autonomous Driving Services

SafetyDGX agent

arXiv:2606.07067v1 Announce Type: new Abstract: Safety is a fundamental requirement in the development of autonomous driving (AD) systems. While function offloading has demonstrated significant benefi

SafeGene: Reusable Adapters for Transferable Safety Alignment

SafetyDGX agent

arXiv:2606.06519v1 Announce Type: new Abstract: Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vul

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

SafetyDGX agent

arXiv:2606.06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework f

5 Jun 2026

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

SafetyDGX agent

arXiv:2606.05743v1 Announce Type: cross Abstract: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifi

29 May 2026

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

SafetyDGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

12 May 2026

Synergistic Simplex: Cooperative Runtime Assurance for Safety-Critical Autonomous Systems

SafetyDGX agent

arXiv:2605.08190v1 Announce Type: new Abstract: Autonomous systems increasingly rely on machine-learning (ML) components for safety-critical tasks such as perception and control in autonomous vehicles

Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning

SafetyDGX agent

arXiv:2602.17546v2 Announce Type: replace Abstract: Instruction-following language models are trained to be helpful and safe, yet their safety behavior can deteriorate under benign fine-tuning and wor

Mental Health AI Safety Claims Must Preserve Temporal Evidence

SafetyDGX agent

arXiv:2605.08827v1 Announce Type: new Abstract: The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, o

6 May 2026

Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation

SafetyDGX agent

arXiv:2603.17368v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning cap

Multilingual Safety Alignment via Self-Distillation

SafetyDGX agent

arXiv:2605.02971v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain hig

1 May 2026

Policy-Grounded Safety Evaluation of 20 Large Language Models

SafetyDGX agent

arXiv:2507.14719v2 Announce Type: replace Abstract: As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. T

28 Apr 2026

Discovering Agentic Safety Specifications from 1-Bit Danger Signals

SafetyDGX agent

arXiv:2604.23210v1 Announce Type: new Abstract: Can large language model agents discover hidden safety objectives through experience alone? We introduce EPO-Safe (Experiential Prompt Optimization for

Grammar-Constrained Refinement of Safety Operational Rules Using Language in the Loop: What Could Go Wrong

SafetyDGX agent

arXiv:2604.23523v1 Announce Type: cross Abstract: Safety specifications in cyber-physical systems (CPS) capture the operational conditions the system must satisfy to operate safely within its intended

27 Apr 2026

When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models

SafetyDGX agent

arXiv:2510.21285v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex multi-step reasoning, yet they still exhibit severe safety failures such as harm

21 Apr 2026

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

Model ReleasesDGX agent

arXiv:2503.03480v4 Announce Type: replace Abstract: Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-w

Continual Safety Alignment via Gradient-Based Sample Selection

SafetyDGX agent

arXiv:2604.17215v1 Announce Type: new Abstract: Large language models require continuous adaptation to new tasks while preserving safety alignment. However, fine-tuning on even benign data often compr

On Safety Risks in Experience-Driven Self-Evolving Agents

SafetyDGX agent

arXiv:2604.16968v1 Announce Type: new Abstract: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self

17 Apr 2026

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

SafetyDGX agent

arXiv:2603.13933v2 Announce Type: replace Abstract: Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on a

RL-STPA: Adapting System-Theoretic Hazard Analysis for Safety-Critical Reinforcement Learning

SafetyDGX agent

arXiv:2604.15201v1 Announce Type: new Abstract: As reinforcement learning (RL) deployments expand into safety-critical domains, existing evaluation methods fail to systematically identify hazards aris

23 Jul 2026

Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards

SafetyDGX agent

arXiv:2607.19827v1 Announce Type: new Abstract: Ensuring safety in Physical AI systems operating in real-world environments is a critical challenge, particularly in hospital wards where vulnerable pat

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

SafetyDGX agent

arXiv:2607.19913v1 Announce Type: new Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-orient

16 Jul 2026

Designing Safety-Constrained LLM Systems for Public Health Information Access

SafetyDGX agent

arXiv:2607.13038v1 Announce Type: cross Abstract: We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on m

4 Jun 2026

When Autoregressive Consistency Hurts Safety Alignment

SafetyDGX agent

arXiv:2606.04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near t

3 Jun 2026

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

SafetyDGX agent

arXiv:2606.03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identifica

2 Jun 2026

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning

SafetyDGX agent

arXiv:2606.00160v1 Announce Type: cross Abstract: Large language models (LLMs) suffer from degraded safety capabilities even when fine-tuned with benign datasets. However, existing methods for identif

26 May 2026

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

SafetyDGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

19 May 2026

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction

SafetyDGX agent

arXiv:2605.18104v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often fail to transfer safety capabilities learned in the text modality to semantically equivalent non-text inp

14 May 2026

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions

SafetyDGX agent

arXiv:2408.12935v4 Announce Type: replace Abstract: AI Safety is an emerging area of critical importance to the safe adoption and deployment of AI systems. With the rapid proliferation of AI and espec

13 May 2026

BSO: Safety Alignment Is Density Ratio Matching

SafetyDGX agent

arXiv:2605.12339v1 Announce Type: new Abstract: Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement lear

5 May 2026

Online Safety Filter for Deformable Object Manipulation with Horizon Agnostic Neural Operators

SafetyDGX agent

arXiv:2605.01069v1 Announce Type: new Abstract: Safety critical control of robotic manipulation tasks involving deformable media such as fluids, cloth, and soft objects remains challenging because exi

24 Apr 2026

SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs

SafetyDGX agent

arXiv:2604.20930v1 Announce Type: cross Abstract: Internal Safety Collapse (ISC) is a failure mode in which frontier LLMs, when executing legitimate professional tasks whose correct completion structu

12 Aug 2026

The Illusion of Cross-Lingual Safety in Low-Resource Languages

Local AiDGX agent

arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. How

11 Aug 2026

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

SafetyDGX agent

arXiv:2608.07535v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding an

Safety Cost of Steering Vectors Is Separable and Reducible

SafetyDGX agent

arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally comprom

31 Jul 2026

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

SafetyDGX agent

arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, a

25 Jun 2026

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion

SafetyDGX agent

arXiv:2606.25097v1 Announce Type: new Abstract: Speculative decoding accelerates inference by letting a draft model propose tokens for a target model to verify, raising a concrete safety question: at

← Previous
1234…238
Next →