AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
21 Apr 2026

RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification

SafetyDGX agent

arXiv:2604.17391v1 Announce Type: cross Abstract: RISC-V is emerging as a viable platform for automotive-grade embedded computing, with recent ISO 26262 ASIL-D certifications demonstrating readiness f

Why Agents Compromise Safety Under Pressure

SafetyDGX agent

arXiv:2603.14975v2 Announce Type: replace-cross Abstract: Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent…

SafetyDGX agent

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

SafetyDGX agent

arXiv:2405.13068v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Cu

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints

SafetyDGX agent

arXiv:2604.16916v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respo

17 Apr 2026

Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety

SafetyDGX agent

arXiv:2604.13101v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into aviation safety decision-making represents a significant technological advancement, yet their sta

12 Aug 2026

Injecting Hallucinations in Autonomous Vehicles: A Component-Agnostic Safety Evaluation Framework

SafetyDGX agent

arXiv:2510.07749v2 Announce Type: replace Abstract: Perception failures in autonomous vehicles (AV) remain a major safety concern because they are the basis for many accidents. To study how these fail

11 Aug 2026

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

Model ReleasesDGX agent

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

10 Aug 2026

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Model ReleasesDGX agent

arXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech

5 Aug 2026

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

Local AiDGX agent

arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward wh

Toward Certified Functional Safety for Industrial Humanoid Robots: The Fail-Passive Gap and a Feasibility Study

SafetyDGX agent

arXiv:2608.02809v1 Announce Type: new Abstract: Industrial humanoid robots are constrained less by locomotion or manipulation capability than by the immaturity of functional safety certification for l

4 Aug 2026

Towards General Language-Conditioned Latent Safety Filters

SafetyDGX agent

arXiv:2608.00315v1 Announce Type: cross Abstract: Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified

28 Jul 2026

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

Model ReleasesDGX agent

arXiv:2607.22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation,

From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

SafetyDGX agent

arXiv:2510.03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastru

27 Jul 2026

Nvidia forms the Open Secure AI Alliance, a coalition including CrowdStrike, Hugging Face, and Dell to develop and share tools for AI safety and cybersecurity (Jaspreet Singh/Reuters)

SafetyDGX agent

Jaspreet Singh / Reuters: Nvidia forms the Open Secure AI Alliance, a coalition including CrowdStrike, Hugging Face, and Dell to develop and share tools for AI safety and cybersecurity — Nvidia (NVDA.

7 Jul 2026

CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation

SafetyDGX agent

arXiv:2607.04451v1 Announce Type: new Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating con

3 Jul 2026

Safety Targeted Embedding Exploit via Refinement

Model ReleasesDGX agent

arXiv:2607.01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-r

2 Jul 2026

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire

SafetyDGX agent

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire Altman’s AI safety proposal: let us win, or everybody loses https://ft.trib.al/UrDI

30 Jun 2026

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

SafetyDGX agent

arXiv:2606.29581v1 Announce Type: cross Abstract: Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluations usually

29 Jun 2026

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

SafetyDGX agent

arXiv:2510.16492v4 Announce Type: replace Abstract: As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While

25 Jun 2026

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

Model ReleasesDGX agent

arXiv:2606.25442v1 Announce Type: new Abstract: Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. Ho

23 Jun 2026

SHIELD: Safety on Humanoids via CBFs In Expectation on Learned Dynamics

SafetyDGX agent

arXiv:2505.11494v3 Announce Type: replace Abstract: Robot learning has produced remarkably effective ``black-box'' controllers for complex tasks such as dynamic locomotion on humanoids. Yet ensuring d

11 Jun 2026

Ideagram 4 - Safety Filter?

SafetyDGX agent

Ideogram 4 includes runtime safety filters powered by Hive for prompt and output moderation , with NSFW prompts blocked by displaying 'Image blocked by safety filter' . Users have reported false posit

10 Jun 2026

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

SafetyDGX agent

arXiv:2606.06037v2 Announce Type: cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on m

9 Jun 2026

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

SafetyDGX agent

arXiv:2606.07678v1 Announce Type: cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data se

8 Jun 2026

Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering

SafetyDGX agent

arXiv:2606.07193v1 Announce Type: new Abstract: Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent duri

Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring

Model ReleasesDGX agent

arXiv:2603.25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs). However, unsafe events are rare in real-world CPS operations, creating an extreme

3 Jun 2026

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

SafetyDGX agent

arXiv:2606.03793v1 Announce Type: new Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

Model ReleasesDGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

2 Jun 2026

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing

SafetyDGX agent

arXiv:2606.00686v1 Announce Type: new Abstract: The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmfu

Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment

SafetyDGX agent

arXiv:2606.00369v1 Announce Type: cross Abstract: Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets rema

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

SafetyDGX agent

arXiv:2606.00611v1 Announce Type: new Abstract: Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moder

1 Jun 2026

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

SafetyDGX agent

arXiv:2605.31381v1 Announce Type: new Abstract: We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that

28 May 2026

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

Model ReleasesDGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

SafetyDGX agent

arXiv:2605.28338v1 Announce Type: new Abstract: Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

SafetyDGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

SafetyDGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

27 May 2026

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

SafetyDGX agent

arXiv:2605.26947v1 Announce Type: new Abstract: Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt datase

26 May 2026

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

SafetyDGX agent

arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align

21 May 2026

Towards Context-Invariant Safety Alignment for Large Language Models

SafetyDGX agent

arXiv:2605.20994v1 Announce Type: new Abstract: Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a stand

19 May 2026

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey

SafetyDGX agent

arXiv:2505.17352v2 Announce Type: replace Abstract: Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignmen

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

Model ReleasesDGX agent

arXiv:2605.16282v1 Announce Type: cross Abstract: The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a prolifera

13 May 2026

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

SafetyDGX agent

arXiv:2602.07892v2 Announce Type: replace-cross Abstract: Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility

12 May 2026

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models

SafetyDGX agent

arXiv:2605.08513v1 Announce Type: cross Abstract: Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expr

11 May 2026

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

Model ReleasesDGX agent

arXiv:2503.06223v5 Announce Type: replace Abstract: Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is

7 May 2026

ChatGPT’s ‘Trusted Contact’ will alert loved ones of safety concerns

SafetyDGX agent

OpenAI is launching an optional safety feature for ChatGPT that allows adult users to assign an emergency contact for mental health and safety concerns. Friends, family members, or caregivers designat

Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation

SafetyDGX agent

arXiv:2605.04366v1 Announce Type: cross Abstract: Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation off

6 May 2026

FORMULA: FORmation MPC with neUral barrier Learning for safety Assurance

SafetyDGX agent

arXiv:2604.04409v2 Announce Type: replace Abstract: Multi-robot systems (MRS) are essential for large-scale applications such as disaster response, material transport, and warehouse logistics, yet ens

5 May 2026

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

Model ReleasesDGX agent

arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable t

30 Apr 2026

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

SafetyDGX agent

arXiv:2604.26577v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

SafetyDGX agent

arXiv:2603.02259v2 Announce Type: replace-cross Abstract: Multi-agent systems provide mature methodologies for role decomposition, coordination, and normative governance, capabilities that remain esse

28 Apr 2026

AI Safety Training Can be Clinically Harmful

SafetyDGX agent

arXiv:2604.23445v1 Announce Type: cross Abstract: Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigo

Our commitment to community safety

SafetyDGX agent

OpenAI outlines its commitment to implementing safety measures and responsible practices in the development and deployment of AI systems to protect users and communities. The statement likely covers O

27 Apr 2026

Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems

SafetyDGX agent

arXiv:2604.22244v1 Announce Type: new Abstract: Ensuring safety for black-box hybrid dynamical systems presents significant challenges due to their instantaneous state jumps and unknown explicit nonli

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem

SafetyDGX agent

arXiv:2506.17299v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their v

24 Apr 2026

Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression

SafetyDGX agent

arXiv:2505.13527v3 Announce Type: replace-cross Abstract: Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to j

23 Apr 2026

A Hough transform approach to safety-aware scalar field mapping using Gaussian Processes

SafetyDGX agent

arXiv:2604.20799v1 Announce Type: new Abstract: This paper presents a framework for mapping unknown scalar fields using a sensor-equipped autonomous robot operating in unsafe environments. The unsafe

LLM-Guided Safety Agent for Edge Robotics with an ISO-Compliant Perception-Compute-Control Architecture

SafetyDGX agent

arXiv:2604.20193v1 Announce Type: new Abstract: Ensuring functional safety in human-robot interaction is challenging because AI perception is inherently probabilistic, whereas industrial standards req

22 Apr 2026

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

SafetyDGX agent

arXiv:2510.08240v2 Announce Type: replace Abstract: Harnessing the power of LLMs requires a delicate dance between being helpful and harmless. This creates a fundamental tension between two competing

19 Apr 2026

Tesla expands its robotaxi service to Dallas and Houston after launching in Austin last year and starting to offer rides without safety drivers in January 2026 (Anthony Ha/TechCrunch)

SafetyDGX agent

Anthony Ha / TechCrunch: Tesla expands its robotaxi service to Dallas and Houston after launching in Austin last year and starting to offer rides without safety drivers in January 2026 — Tesla is expa

← Previous
123456…238
Next →