AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
20 Apr 2026

Trajectory Planning for Safe Dual Control with Active Exploration

SafetyDGX agent

arXiv:2604.15507v1 Announce Type: new Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations, ye

17 Apr 2026

Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades

SafetyDGX agent

arXiv:2604.14251v1 Announce Type: new Abstract: Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalat

16 Apr 2026

I have found that asking for a sestina regularly triggers Opus 4.7's safety guardrails. The forbidden poetic form!

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

Ethan Mollick reported that requesting Claude Opus 4.7 to write sestinas—a complex poetic form with strict structural requirements—frequently triggers the model's safety guardrails, suggesting the AI

Safe and Nonconservative Contingency Planning for Autonomous Vehicles via Online Learning-Based Reachable Set Barriers

SafetyDGX agent

arXiv:2509.07464v2 Announce Type: replace Abstract: Autonomous vehicles must navigate dynamically uncertain environments while balancing safety and efficiency. This challenge is exacerbated by unpredi

15 Apr 2026

ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance

SafetyDGX agent

arXiv:2604.12308v1 Announce Type: new Abstract: Individuals' concerns about data privacy and AI safety are highly contextualized and extend beyond sensitive patterns. Addressing these issues requires

LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

Model ReleasesDGX agent

arXiv:2604.12710v1 Announce Type: cross Abstract: Large language models (LLMs) often demonstrate strong safety performance in high-resource languages, yet exhibit severe vulnerabilities when queried i

14 Apr 2026

LLM-based Realistic Safety-Critical Driving Video Generation

SafetyDGX agent

arXiv:2507.01264v2 Announce Type: replace-cross Abstract: Designing diverse and safety-critical driving scenarios is essential for evaluating autonomous driving systems. In this paper, we propose a no

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

SafetyDGX agent

arXiv:2602.03402v3 Announce Type: replace Abstract: Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerabl

Safety Guarantees in Zero-Shot Reinforcement Learning for Cascade Dynamical Systems

SafetyDGX agent

arXiv:2604.10429v1 Announce Type: new Abstract: This paper considers the problem of zero-shot safety guarantees for cascade dynamical systems. These are systems where a subset of the states (the inner

13 Apr 2026

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

Model ReleasesDGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

8 Apr 2026

Introducing the Child Safety Blueprint

SafetyDGX agent

OpenAI's Child Safety Blueprint, released in April 2026, is a policy framework aimed at combating the rise of AI-enabled child sexual exploitation by combining legal, operational, and technical app...

12 Aug 2026

The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software

SafetyDGX agent

arXiv:2608.10025v1 Announce Type: new Abstract: For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure events experienced by the software) can

11 Aug 2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Model ReleasesDGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

SafetyDGX agent

arXiv:2608.08471v1 Announce Type: new Abstract: Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed

7 Aug 2026

Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics

SafetyDGX agent

arXiv:2608.05656v1 Announce Type: cross Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks fa

4 Aug 2026

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs

Model ReleasesDGX agent

arXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho

From Vessel Trajectories to Safety-Critical Encounter Scenarios: A Generative AI Framework for Autonomous Ship Digital Testing

SafetyDGX agent

arXiv:2603.28067v2 Announce Type: replace Abstract: Digital testing has emerged as a key paradigm for the development and verification of autonomous maritime navigation systems, yet the availability o

3 Aug 2026

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

SafetyDGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

30 Jul 2026

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

SafetyDGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

29 Jul 2026

Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture

SafetyDGX agent

arXiv:2607.24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can b

28 Jul 2026

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Model ReleasesDGX agent

arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, re

24 Jul 2026

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

SafetyDGX agent

arXiv:2607.21151v1 Announce Type: new Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuiti

23 Jul 2026

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

SafetyDGX agent

arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure

it’s just for your safety

SafetyDGX agent

On July 23, 2026, a Polymarket tweet announced that Anthropic donated $20 million to a political nonprofit calling for stricter AI regulation ahead of the U.S. midterm elections. The post, posted at 4

16 Jul 2026

PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner

SafetyDGX agent

arXiv:2603.10330v2 Announce Type: replace-cross Abstract: Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that lear

10 Jul 2026

Efficient Partitioning Method of Large-Scale Public Safety Spatio-Temporal Data based on Information Loss Constraints

SafetyDGX agent

arXiv:2306.12857v3 Announce Type: replace Abstract: The storage, management, and application of massive spatio-temporal data are widely used in practical scenarios, including public safety. However, d

6 Jul 2026

Understanding Annotator Safety Policy with Interpretability

SafetyDGX agent

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such

2 Jul 2026

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

SafetyDGX agent

arXiv:2607.00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed a

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

SafetyDGX agent

arXiv:2607.00334v1 Announce Type: new Abstract: Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous

1 Jul 2026

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

Model ReleasesDGX agent

arXiv:2603.29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing ben

30 Jun 2026

Agent Safety Is Action Alignment

SafetyDGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

Model ReleasesDGX agent

arXiv:2606.28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains p

25 Jun 2026

RAS: Measuring LLM Safety Through Refusal Alignment

Model ReleasesDGX agent

arXiv:2606.25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their

24 Jun 2026

AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

SafetyDGX agent

arXiv:2606.24245v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, th

22 Jun 2026

Nvidia introduces Halos for Robotics to bridge the physical AI safety gap

SafetyDGX agent

Nivida Corp. today announced Halos for Robotics, the industry’s first full framework for robotic safety systems that encompasses building, testing and managing complete artificial intelligence robotic

10 Jun 2026

Canada introduces the Safe Social Media Act, a bill that would ban social media for children under 16 and establish safety standards for AI chatbots (Maria Cheng/Reuters)

SafetyDGX agent

Maria Cheng / Reuters: Canada introduces the Safe Social Media Act, a bill that would ban social media for children under 16 and establish safety standards for AI chatbots — The Canadian government in

Quality Is Not a Safety Proxy Under Quantization

Model ReleasesDGX agent

arXiv:2606.10154v1 Announce Type: new Abstract: Quantized checkpoints are often screened first with quality metrics and only later, if at all, with direct safety tests. This paper audits that shortcut

9 Jun 2026

Impedance MPC for Physical Human-Robot Interaction: Predictive Disturbance Rejection with Joint-Limit Safety

SafetyDGX agent

arXiv:2606.08281v1 Announce Type: new Abstract: Physical human-robot interaction (pHRI) demands simultaneous trajectory accuracy and compliant safety under unplanned contact. Classical impedance contr

Major improvement in safety statistics with FSD turned on in the Netherlands!

SafetyDGX agent

Elon Musk posted on X claiming that Tesla's Full Self-Driving (FSD) system resulted in significant improvements to vehicle safety statistics in the Netherlands. The post suggests FSD activation correl

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation

SafetyDGX agent

arXiv:2606.08414v1 Announce Type: cross Abstract: Diffusion policies have achieved remarkable success in robotic manipulation, yet they often fail to satisfy strict physical constraints required for s

8 Jun 2026

Building Pakistan Notice Helper: A Small AI Tool for a Very Local Safety Problem

SafetyDGX agent

Building Pakistan Notice Helper is a small AI tool designed to address a localized safety issue in Pakistan by helping users understand and process building-related notices. Developed as a hackathon p

6 Jun 2026

Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety

SafetyDGX agent

arXiv:2606.05461v1 Announce Type: new Abstract: Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantifi

5 Jun 2026

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

Model ReleasesDGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

2 Jun 2026

Meta expands Teen Accounts safety features to limit harmful content on Instagram, Facebook, and Messenger, including on nutrition, weight lifting, and anxiety (Eli Tan/New York Times)

SafetyDGX agent

Eli Tan / New York Times: Meta expands Teen Accounts safety features to limit harmful content on Instagram, Facebook, and Messenger, including on nutrition, weight lifting, and anxiety — The changes,

1 Jun 2026

Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

SafetyDGX agent

arXiv:2605.30368v1 Announce Type: cross Abstract: Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SS

29 May 2026

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

Model ReleasesDGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

Model ReleasesDGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

28 May 2026

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

SafetyDGX agent

arXiv:2605.28664v1 Announce Type: cross Abstract: Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are sca

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

Model ReleasesDGX agent

arXiv:2601.01627v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety

Safety-Critical Adaptive Impedance Control via Nonsmooth Control Barrier Functions under State and Input Constraints

SafetyDGX agent

arXiv:2605.28367v1 Announce Type: new Abstract: Safe physical interaction is critical for deploying robotic manipulators in human-robot interaction and contact-rich tasks, where uncertainty, external

27 May 2026

CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment

SafetyDGX agent

arXiv:2603.07211v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes tra

Constrained Meta Reinforcement Learning with Provable Test-Time Safety

SafetyDGX agent

arXiv:2601.21845v2 Announce Type: replace Abstract: Meta reinforcement learning (RL) allows agents to leverage experience across a distribution of tasks on which the agent can train at will, enabling

Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection

SafetyDGX agent

arXiv:2605.27155v1 Announce Type: cross Abstract: Testing object detectors in safety-critical domains requires semantically meaningful probes beyond pixel-level corruptions. We present SemProbe, a too

25 May 2026

CarlaNCAP: A Framework for Quantifying the Safety of Vulnerable Road Users in Infrastructure-Assisted Collective Perception Using EuroNCAP Scenarios

SafetyDGX agent

arXiv:2512.11551v2 Announce Type: replace Abstract: The growing number of road users has significantly increased the risk of accidents in recent years. Vulnerable Road Users (VRUs) are particularly at

22 May 2026

Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries

Local AiDGX agent

arXiv:2605.21712v1 Announce Type: new Abstract: Transportation safety analysis requires integrating crash records, roadway attributes, and geospatial data through GIS-based workflows, but access remai

21 May 2026

Anomaly-Informed Confidence Calibration for Vision-Based Safety Prediction

SafetyDGX agent

arXiv:2605.21109v1 Announce Type: new Abstract: Reliable confidence estimates are important for safely deploying vision-based controllers in autonomous racing, where safety predictions must be derived

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry

Model ReleasesDGX agent

arXiv:2605.20241v1 Announce Type: cross Abstract: Prompt-level safety probes for large language models use hidden-state representations to separate safe from unsafe prompts, but strong average detecti

Time-To-Reach Separation and Safety Filtering for Safe, Fair, and Efficient Multi-Agent Coordination

SafetyDGX agent

arXiv:2605.20625v1 Announce Type: cross Abstract: Advanced Air Mobility (AAM) operations are expected to significantly increase aerial traffic in urban airspace, requiring autonomous traffic managemen

19 May 2026

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

SafetyDGX agent

arXiv:2605.17413v1 Announce Type: cross Abstract: Safety-aligned language models often refuse cybersecurity requests whose wording resembles misuse, even when the task is authorized and defensive. Thi

LLM-Safety Evaluations Lack Robustness

SafetyDGX agent

arXiv:2503.02574v2 Announce Type: replace-cross Abstract: In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of n

← Previous
1…45678…238
Next →