AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
19 May 2026

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

SafetyDGX agent

arXiv:2603.18178v2 Announce Type: replace-cross Abstract: The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-co

18 May 2026

IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia

Model ReleasesDGX agent

arXiv:2603.17915v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are deployed in multilingual settings, their safety behavior in culturally diverse, low-resource languages rem

15 May 2026

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2509.26100v2 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, exis

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

SafetyDGX agent

arXiv:2605.15164v1 Announce Type: cross Abstract: This position paper argues that behavioural assurance, even when carefully designed, is being asked to carry safety claims it cannot verify. AI govern

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

Model ReleasesDGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

14 May 2026

Before the Last Token: Diagnosing Final-Token Safety Probe Failures

SafetyDGX agent

arXiv:2605.12726v1 Announce Type: new Abstract: Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompts can contain probe-visible unsafe evidence distribute

13 May 2026

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

SafetyDGX agent

arXiv:2602.07668v2 Announce Type: replace Abstract: The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state

Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

SafetyDGX agent

arXiv:2605.11769v1 Announce Type: new Abstract: Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. Whi

12 May 2026

Containment Verification: AI Safety Guarantees Independent of Alignment

SafetyDGX agent

arXiv:2605.09045v1 Announce Type: new Abstract: Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain

6 May 2026

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

SafetyDGX agent

arXiv:2605.02900v1 Announce Type: cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, saf

5 May 2026

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection

SafetyDGX agent

arXiv:2509.00673v2 Announce Type: replace Abstract: We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety

TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks

SafetyDGX agent

arXiv:2605.01761v1 Announce Type: new Abstract: Text-to-Video (T2V) models have demonstrated remarkable capability in generating temporally coherent videos from natural language prompts, yet they also

30 Apr 2026

A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety

SafetyDGX agent

arXiv:2505.17342v2 Announce Type: replace Abstract: Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and

From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Model

SafetyDGX agent

arXiv:2604.26052v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes such as attack success rate, refusal rate, or harmful/not-harmful re

24 Apr 2026

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

Model ReleasesDGX agent

arXiv:2601.18491v2 Announce Type: replace Abstract: The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current gua

22 Apr 2026

Reasoning Structure Matters for Safety Alignment of Reasoning Models

SafetyDGX agent

arXiv:2604.18946v1 Announce Type: new Abstract: Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This

Vision-Based Human Awareness Estimation for Enhanced Safety and Efficiency of AMRs in Industrial Warehouses

SafetyDGX agent

arXiv:2604.18627v1 Announce Type: new Abstract: Ensuring human safety is of paramount importance in warehouse environments that feature mixed traffic of human workers and autonomous mobile robots (AMR

When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains

Model ReleasesDGX agent

arXiv:2604.19001v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm

21 Apr 2026

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

SafetyDGX agent

arXiv:2604.16424v1 Announce Type: cross Abstract: State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

Model ReleasesDGX agent

arXiv:2601.04043v2 Announce Type: replace Abstract: As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to

15 Apr 2026

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

Model ReleasesDGX agent

arXiv:2509.18127v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, un

13 Apr 2026

Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies

SafetyDGX agent

arXiv:2604.09189v1 Announce Type: cross Abstract: LLMs internalize safety policies through RLHF, yet these policies are never formally specified and remain difficult to inspect. Existing benchmarks ev

OpenKedge: Governing Agentic Mutation with Execution-Bound Safety and Evidence Chains

SafetyDGX agent

arXiv:2604.08601v1 Announce Type: new Abstract: The rise of autonomous AI agents exposes a fundamental flaw in API-centric architectures: probabilistic systems directly execute state mutations without

12 Aug 2026

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

SafetyDGX agent

arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

SafetyDGX agent

arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surroun

Robust Safety Filtering for Input-Constrained Underactuated Linear Systems

SafetyDGX agent

arXiv:2608.10872v1 Announce Type: cross Abstract: We present a robust safety-filtering framework for input-constrained underactuated linear systems subject to unknown disturbances. A baseline H-infty

7 Aug 2026

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

SafetyDGX agent

arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep est

6 Aug 2026

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

SafetyDGX agent

arXiv:2608.04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research ha

SCOPE: Field-of-View-Aware Path Planning in Unknown 3D Environments via Safety-Volume Certification

SafetyDGX agent

arXiv:2608.04420v1 Announce Type: new Abstract: Safe navigation with a body-mounted limited-field-of-view sensor requires the complete robot-inflated volume of an intended motion to be observed and ve

5 Aug 2026

ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories

SafetyDGX agent

arXiv:2608.03866v1 Announce Type: new Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

SafetyDGX agent

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

29 Jul 2026

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

SafetyDGX agent

arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language

28 Jul 2026

Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety

SafetyDGX agent

arXiv:2607.24168v1 Announce Type: new Abstract: Road crashes remain among the gravest threats to public safety, and preventing them is a defining task of transportation systems worldwide. Much of that

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

SafetyDGX agent

arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whe

16 Jul 2026

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

SafetyDGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

10 Jul 2026

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

Model ReleasesDGX agent

arXiv:2607.08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task

Efficient Safety Alignment of Language Models via Latent Personality Traits

SafetyDGX agent

arXiv:2607.07918v1 Announce Type: cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Late

7 Jul 2026

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on a…

SafetyDGX agent

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on as an afterthought. Because trust is everything when robots s

VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

SafetyDGX agent

arXiv:2607.05180v1 Announce Type: cross Abstract: Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the ve

6 Jul 2026

Q&A with Agility Robotics CEO Peggy Johnson on why the startup is going public via SPAC, the physical layer as its proprietary advantage, safety, and more (Connie Loizos/TechCrunch)

SafetyDGX agent

Connie Loizos / TechCrunch: Q&A with Agility Robotics CEO Peggy Johnson on why the startup is going public via SPAC, the physical layer as its proprietary advantage, safety, and more — The humanoid ro

5 Jul 2026

Tesla rolls out its Robotaxi service without a safety monitor in Miami, its fifth city, as it aims to expand to a dozen US states by the end of 2026 (Grace Kay/The Information)

SafetyDGX agent

Grace Kay / The Information: Tesla rolls out its Robotaxi service without a safety monitor in Miami, its fifth city, as it aims to expand to a dozen US states by the end of 2026 — Tesla said it rolled

3 Jul 2026

Multi-modal Rail Crossing Safety Analysis

SafetyDGX agent

arXiv:2607.01365v1 Announce Type: cross Abstract: Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our abil

2 Jul 2026

FastBridge: Closing the Model-Based Realization Gap in Safety Filters on 3D Gaussian Splatting for Fast Quadrotor Flight

SafetyDGX agent

arXiv:2607.01200v1 Announce Type: new Abstract: Fast quadrotor flight requires safe obstacle avoidance under tight onboard compute limits. While 3D Gaussian Splatting (3DGS) provides a continuous, geo

1 Jul 2026

Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios

Model ReleasesDGX agent

arXiv:2601.21173v2 Announce Type: replace-cross Abstract: With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in com

30 Jun 2026

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

Model ReleasesDGX agent

arXiv:2606.30256v1 Announce Type: new Abstract: Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides p

29 Jun 2026

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

Model ReleasesDGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now s…

SafetyDGX agent

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now so it’s taking things slow and smoothly. Highway performance

26 Jun 2026

Real-Time Safety Evaluation of Human Arm Operations Using a Wrist-Mounted IMU with PSM System

Model ReleasesDGX agent

arXiv:2502.09241v2 Announce Type: replace Abstract: This paper presents a novel approach to real-time safety monitoring in human-robot collaborative manufacturing environments through a wrist-mounted

25 Jun 2026

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know wh…

SafetyDGX agent

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know what risks everyone will face if/when open source reaches Myth

24 Jun 2026

Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications

SafetyDGX agent

arXiv:2606.23858v1 Announce Type: cross Abstract: A primary challenge in AI safety is the existence of adversarial examples -- slightly distorted inputs that cause a neural network (NN) to misclassify

23 Jun 2026

IndicGuard: A Multilingual Safety Guard Model and Dataset for Indic Languages

Model ReleasesDGX agent

arXiv:2606.22841v1 Announce Type: cross Abstract: As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes, ensuring their safety and alignment with regional

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

Model ReleasesDGX agent

arXiv:2606.22686v1 Announce Type: cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work, we investig

22 Jun 2026

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alig…

SafetyDGX agent

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alignment. You can read it here: https://arxiv.org/abs/2606.1691

Nvidia unveils Halos, a safety-focused OS developed from autonomous vehicle tech and designed to run on IGX Thor hardware for humanoid robots, and opens a lab (Ian King/Bloomberg)

SafetyDGX agent

Ian King / Bloomberg: Nvidia unveils Halos, a safety-focused OS developed from autonomous vehicle tech and designed to run on IGX Thor hardware for humanoid robots, and opens a lab — Nvidia Corp. is w

11 Jun 2026

Benchmarking Large Language Models for Safety Data Extraction

Model ReleasesDGX agent

arXiv:2606.11204v1 Announce Type: new Abstract: Accurate extraction of structured information from Safety Data Sheets (SDS) remains challenging in industrial safety due to heterogeneous document forma

Schutzen: Evaluating LLM Safety in Bulgarian and German Contexts

SafetyDGX agent

arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disr

10 Jun 2026

An essay on policy responses to AI's exponential progress across regulation and public safety, macroeconomics and taxes, science, civil liberties, geopolitics (Dario Amodei)

SafetyDGX agent

Dario Amodei: An essay on policy responses to AI's exponential progress across regulation and public safety, macroeconomics and taxes, science, civil liberties, geopolitics — In one of the side plots

CameraMatics, which uses AI to help fleet operators improve safety, reduce operational risk, and lower carbon emissions, raised €49M (Joe Brennan/The Irish Times)

SafetyDGX agent

Joe Brennan / The Irish Times: CameraMatics, which uses AI to help fleet operators improve safety, reduce operational risk, and lower carbon emissions, raised €49M — Tech focuses on accident preventio

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a…

SafetyDGX agent

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a different kind of answer about what is real and what is mark

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

SafetyDGX agent

arXiv:2606.09866v1 Announce Type: cross Abstract: Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods

← Previous
1…56789…238
Next →