AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
16 Apr 2026

Hierarchical Reinforcement Learning with Runtime Safety Shielding for Power Grid Operation

Model ReleasesDGX agent

arXiv:2604.14032v1 Announce Type: cross Abstract: Reinforcement learning has shown promise for automating power-grid operation tasks such as topology control and congestion management. However, its de

15 Apr 2026

Goal-Conditioned Neural ODEs with Guaranteed Safety and Stability for Learning-Based All-Pairs Motion Planning

SafetyDGX agent

arXiv:2604.02821v2 Announce Type: replace Abstract: This paper presents a learning-based approach for all-pairs motion planning, where the initial and goal states are allowed to be arbitrary points in

12 Aug 2026

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2607.16943v2 Announce Type: replace Abstract: Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenari

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

Model ReleasesDGX agent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do

11 Aug 2026

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

SafetyDGX agent

arXiv:2608.08485v1 Announce Type: new Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inf

10 Aug 2026

Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems

SafetyDGX agent

arXiv:2608.06378v1 Announce Type: cross Abstract: Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly focus on drive

7 Aug 2026

JTA: Joint Testability Architecture for Scenario-Based Validation of Safety-Critical Software

SafetyDGX agent

arXiv:2608.05594v1 Announce Type: cross Abstract: Validation adequacy in safety-critical software depends on more than the system under test. Critical scenarios must be constructed under controlled co

6 Aug 2026

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

SafetyDGX agent

arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific

5 Aug 2026

S^3: Improving Agent Safety through Multi-Stage Defense

Model ReleasesDGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

3 Aug 2026

MROPE: A Multi-Robot Safe Cooperative Strategy via combined Predictive Safety Filters and Ellipse-based Constraint Compression

SafetyDGX agent

arXiv:2607.29203v1 Announce Type: new Abstract: Deploying drone swarms to track a dynamic target in cluttered environments presents severe computational and safety challenges. We propose MROPE, a hier

31 Jul 2026

Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning

SafetyDGX agent

arXiv:2607.27626v1 Announce Type: new Abstract: Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor's peak Age of Inform

29 Jul 2026

Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling

SafetyDGX agent

arXiv:2603.14841v3 Announce Type: replace-cross Abstract: Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offer l

28 Jul 2026

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Model ReleasesDGX agent

arXiv:2607.15166v2 Announce Type: replace Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? W

23 Jul 2026

Matching Ranks Over Probability Yields Truly Deep Safety Alignment

SafetyDGX agent

arXiv:2512.05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their 'open' nature introduces more avenues fo

20 Jul 2026

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.…

SafetyDGX agent

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running

16 Jul 2026

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

SafetyDGX agent

arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature

15 Jul 2026

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

SafetyDGX agent

arXiv:2607.12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequentl

7 Jul 2026

Detecting Architectural Drift in Safety-Critical Firmware through Runtime Trace Analysis

SafetyDGX agent

arXiv:2607.03135v1 Announce Type: cross Abstract: Maintaining consistency between architectural design and runtime-observed behavior is challenging in long-lived safety-critical firmware. This paper p

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

Model ReleasesDGX agent

arXiv:2606.20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under

30 Jun 2026

Agentic Safety is an Epistemic Property, Not a Behavioral One

SafetyDGX agent

arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods

CAREBench: A Child-Safety Risk Benchmark for Language Models

Model ReleasesDGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

29 Jun 2026

Algorithms for Deciding the Safety of States in Fully Observable Non-deterministic Problems: Technical Report

SafetyDGX agent

arXiv:2603.15282v2 Announce Type: replace Abstract: Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees. Recent work introduced

25 Jun 2026

Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations

SafetyDGX agent

arXiv:2606.25629v1 Announce Type: new Abstract: Robot navigation in safety-critical scenarios faces significant challenges from unforeseen semantic events, where collisions arise primarily from the un

24 Jun 2026

Verifiable Foundation Models for Robot Safety

SafetyDGX agent

arXiv:2606.23754v1 Announce Type: cross Abstract: Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes the

9 Jun 2026

How Well Do Latent World Models Understand Partially Observable Safety Constraints?

SafetyDGX agent

arXiv:2510.06492v2 Announce Type: replace Abstract: Latent world models are a promising approach for learning state representations and dynamics directly from high-dimensional observations, enabling r

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

SafetyDGX agent

arXiv:2606.09165v1 Announce Type: new Abstract: Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle u

Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models

SafetyDGX agent

arXiv:2606.08451v1 Announce Type: cross Abstract: Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy. Althoug

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

SafetyDGX agent

arXiv:2606.08531v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, a

5 Jun 2026

Learning of Robot Safety Policies via Adversarial Synthetic Scenarios

SafetyDGX agent

arXiv:2606.05952v1 Announce Type: new Abstract: In this work, we propose an agentic gamification framework for hazard-informed learning of robot safety policies through synthetic scenarios. We model s

RiskFlow: Fast and Faithful Safety-Critical Traffic Scenario Generation

SafetyDGX agent

arXiv:2606.06423v1 Announce Type: new Abstract: Safety-critical traffic scenario generation is essential for evaluating autonomous driving systems under rare but high-risk interactions. Existing diffu

3 Jun 2026

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

Model ReleasesDGX agent

arXiv:2606.03954v1 Announce Type: new Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible conse

2 Jun 2026

MESA: Improving MoE Safety Alignment via Decentralized Expertise

SafetyDGX agent

arXiv:2606.00651v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures scale Large Language Models (LLMs) efficiently, enabling greater capacity with reduced computational cost by dy

Safety Alignment of LMs via Non-cooperative Games

SafetyDGX agent

arXiv:2512.20806v3 Announce Type: replace Abstract: Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rel

29 May 2026

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

Model ReleasesDGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

Model ReleasesDGX agent

arXiv:2605.28830v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a c

28 May 2026

SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving

SafetyDGX agent

arXiv:2605.28583v1 Announce Type: cross Abstract: Ensuring both safety and efficiency in decision-making for autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement

27 May 2026

Curriculum Learning for Safety Alignment

SafetyDGX agent

arXiv:2605.26315v1 Announce Type: cross Abstract: Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibit

21 May 2026

AI-based Prediction of Independent Construction Safety Outcomes from Universal Attributes

SafetyDGX agent

arXiv:1908.05972v3 Announce Type: replace Abstract: This paper significantly improves on, and finishes to validate, an approach proposed in previous research in which safety outcomes were predicted fr

PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment

SafetyDGX agent

arXiv:2605.21225v1 Announce Type: new Abstract: We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it f

20 May 2026

Distributional AGI Safety

SafetyDGX agent

arXiv:2512.16856v2 Announce Type: replace Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an e

18 May 2026

Reactive Robot-Centric Safety for Autonomous Navigation in Constrained and Dynamic Environments

SafetyDGX agent

arXiv:2605.15782v1 Announce Type: new Abstract: In this work, we address the problem of ensuring real-time safety in autonomous robot navigation, in spatially constrained dynamic environments, by util

13 May 2026

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

Model ReleasesDGX agent

arXiv:2605.12386v1 Announce Type: new Abstract: Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are tempo

12 May 2026

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

SafetyDGX agent

arXiv:2601.18061v3 Announce Type: replace Abstract: Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI

GuardAD: Safeguarding Autonomous Driving MLLMs via Markovian Safety Logic

SafetyDGX agent

arXiv:2605.10386v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse sa

11 May 2026

Why Does Agentic Safety Fail to Generalize Across Tasks?

SafetyDGX agent

arXiv:2605.06992v1 Announce Type: new Abstract: AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen

7 May 2026

Safety by Invariance, Liveness through Refinement: Heterogeneous Contract Framework for Co-Design of Layered Control

SafetyDGX agent

arXiv:2605.04222v1 Announce Type: cross Abstract: Real-world control systems must achieve long-horizon objectives (liveness) while respecting continuous-time safety constraints, a combination that mot

6 May 2026

Safety-critical Control Under Partial Observability: Reach-Avoid POMDP meets Belief Space Control

SafetyDGX agent

arXiv:2603.10572v2 Announce Type: replace Abstract: Partially Observable Markov Decision Processes (POMDPs) provide a principled framework for robot decision-making under uncertainty. Solving reach-av

4 May 2026

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

SafetyDGX agent

arXiv:2605.00227v1 Announce Type: new Abstract: There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely

Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification

SafetyDGX agent

arXiv:2605.00326v1 Announce Type: new Abstract: Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they a

1 May 2026

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

SafetyDGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

Focus Session: Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification

SafetyDGX agent

arXiv:2604.27807v1 Announce Type: new Abstract: The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

SafetyDGX agent

arXiv:2604.28087v1 Announce Type: cross Abstract: Rule-based systems remain central in safety-critical domains but often struggle with scalability, brittleness, and goal misspecification. These limita

30 Apr 2026

A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

SafetyDGX agent

arXiv:2505.10924v4 Announce Type: replace-cross Abstract: Recently, AI-driven interactions with computing devices have advanced from basic prototype tools to sophisticated, LLM-based systems that emul

Recipes for Calibration Checks in Safety-Critical Applications

SafetyDGX agent

arXiv:2604.26479v1 Announce Type: cross Abstract: Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters

27 Apr 2026

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance t…

SafetyDGX agent

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance to adoption. In this op-ed in the @FT, I underscore that Euro

23 Apr 2026

US child safety group NCMEC received 1.5M reports of suspected CSAM with ties to AI in 2025, a significant surge compared to 67,000 in 2024 and 4,700 in 2023 (Bloomberg)

SafetyDGX agent

Bloomberg: US child safety group NCMEC received 1.5M reports of suspected CSAM with ties to AI in 2025, a significant surge compared to 67,000 in 2024 and 4,700 in 2023 — William Michael Haslach was a

22 Apr 2026

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

SafetyDGX agent

arXiv:2509.26238v4 Announce Type: replace Abstract: Monitoring large language models' (LLMs) activations is an effective way to detect harmful requests before they lead to unsafe outputs. However, tra

21 Apr 2026

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

Model ReleasesDGX agent

arXiv:2505.15404v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enha

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

Model ReleasesDGX agent

arXiv:2604.17691v1 Announce Type: new Abstract: Safety alignment in large language models is remarkably shallow: it is concentrated in the first few output tokens and reversible by fine-tuning on as f

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

Local AiDGX agent

arXiv:2604.16542v1 Announce Type: cross Abstract: Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). How

← Previous
1…34567…238
Next →