AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Model Releases

Hierarchical Reinforcement Learning with Runtime Safety Shielding for Power Grid Operation

DGX agent

arXiv:2604.14032v1 Announce Type: cross Abstract: Reinforcement learning has shown promise for automating power-grid operation tasks such as topology control and congestion management. However, its de

model-releasesarxiv-cs-lg
16 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Goal-Conditioned Neural ODEs with Guaranteed Safety and Stability for Learning-Based All-Pairs Motion Planning

DGX agent

arXiv:2604.02821v2 Announce Type: replace Abstract: This paper presents a learning-based approach for all-pairs motion planning, where the initial and goal states are allowed to be arbitrary points in

safetyarxiv-cs-ro
15 Apr 2026
Model Releases

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

DGX agent

arXiv:2607.16943v2 Announce Type: replace Abstract: Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenari

model-releasesarxiv-cs-ro
12 Aug 2026
Model Releases

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

DGX agent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

DGX agent

arXiv:2608.08485v1 Announce Type: new Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inf

safetyarxiv-cs-ai
11 Aug 2026
Safety

Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems

DGX agent

arXiv:2608.06378v1 Announce Type: cross Abstract: Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly focus on drive

safetyarxiv-cs-ai
10 Aug 2026
Safety

JTA: Joint Testability Architecture for Scenario-Based Validation of Safety-Critical Software

DGX agent

arXiv:2608.05594v1 Announce Type: cross Abstract: Validation adequacy in safety-critical software depends on more than the system under test. Critical scenarios must be constructed under controlled co

safetyarxiv-cs-ro
7 Aug 2026
Safety

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

DGX agent

arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific

safetyarxiv-cs-cl
6 Aug 2026
Model Releases

S^3: Improving Agent Safety through Multi-Stage Defense

DGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

MROPE: A Multi-Robot Safe Cooperative Strategy via combined Predictive Safety Filters and Ellipse-based Constraint Compression

DGX agent

arXiv:2607.29203v1 Announce Type: new Abstract: Deploying drone swarms to track a dynamic target in cluttered environments presents severe computational and safety challenges. We propose MROPE, a hier

safetyarxiv-cs-ro
3 Aug 2026
Safety

Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning

DGX agent

arXiv:2607.27626v1 Announce Type: new Abstract: Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor's peak Age of Inform

safetyarxiv-cs-lg
31 Jul 2026
Safety

Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling

DGX agent

arXiv:2603.14841v3 Announce Type: replace-cross Abstract: Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offer l

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

DGX agent

arXiv:2607.15166v2 Announce Type: replace Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? W

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Matching Ranks Over Probability Yields Truly Deep Safety Alignment

DGX agent

arXiv:2512.05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their 'open' nature introduces more avenues fo

safetyarxiv-cs-ai
23 Jul 2026
Safety

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.…

DGX agent

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running

safetyopenai--x
20 Jul 2026
Safety

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

DGX agent

arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature

safetyarxiv-cs-ai
16 Jul 2026
Safety

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

DGX agent

arXiv:2607.12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequentl

safetyarxiv-cs-ai
15 Jul 2026
Safety

Detecting Architectural Drift in Safety-Critical Firmware through Runtime Trace Analysis

DGX agent

arXiv:2607.03135v1 Announce Type: cross Abstract: Maintaining consistency between architectural design and runtime-observed behavior is challenging in long-lived safety-critical firmware. This paper p

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

DGX agent

arXiv:2606.20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Agentic Safety is an Epistemic Property, Not a Behavioral One

DGX agent

arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

CAREBench: A Child-Safety Risk Benchmark for Language Models

DGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

model-releasesarxiv-cs-lg
30 Jun 2026
Safety

Algorithms for Deciding the Safety of States in Fully Observable Non-deterministic Problems: Technical Report

DGX agent

arXiv:2603.15282v2 Announce Type: replace Abstract: Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees. Recent work introduced

safetyarxiv-cs-ai
29 Jun 2026
Safety

Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations

DGX agent

arXiv:2606.25629v1 Announce Type: new Abstract: Robot navigation in safety-critical scenarios faces significant challenges from unforeseen semantic events, where collisions arise primarily from the un

safetyarxiv-cs-ro
25 Jun 2026
Safety

Verifiable Foundation Models for Robot Safety

DGX agent

arXiv:2606.23754v1 Announce Type: cross Abstract: Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes the

safetyarxiv-cs-lg
24 Jun 2026
Safety

How Well Do Latent World Models Understand Partially Observable Safety Constraints?

DGX agent

arXiv:2510.06492v2 Announce Type: replace Abstract: Latent world models are a promising approach for learning state representations and dynamics directly from high-dimensional observations, enabling r

safetyarxiv-cs-ro
9 Jun 2026
Safety

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

DGX agent

arXiv:2606.09165v1 Announce Type: new Abstract: Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle u

safetyarxiv-cs-ai
9 Jun 2026
Safety

Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models

DGX agent

arXiv:2606.08451v1 Announce Type: cross Abstract: Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy. Althoug

safetyarxiv-cs-ai
9 Jun 2026
Safety

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

DGX agent

arXiv:2606.08531v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, a

safetyarxiv-cs-ai
9 Jun 2026
Safety

Learning of Robot Safety Policies via Adversarial Synthetic Scenarios

DGX agent

arXiv:2606.05952v1 Announce Type: new Abstract: In this work, we propose an agentic gamification framework for hazard-informed learning of robot safety policies through synthetic scenarios. We model s

safetyarxiv-cs-ro
5 Jun 2026
Safety

RiskFlow: Fast and Faithful Safety-Critical Traffic Scenario Generation

DGX agent

arXiv:2606.06423v1 Announce Type: new Abstract: Safety-critical traffic scenario generation is essential for evaluating autonomous driving systems under rare but high-risk interactions. Existing diffu

safetyarxiv-cs-ro
5 Jun 2026
Model Releases

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

DGX agent

arXiv:2606.03954v1 Announce Type: new Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible conse

model-releasesarxiv-cs-cv
3 Jun 2026
Safety

MESA: Improving MoE Safety Alignment via Decentralized Expertise

DGX agent

arXiv:2606.00651v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures scale Large Language Models (LLMs) efficiently, enabling greater capacity with reduced computational cost by dy

safetyarxiv-cs-ai
2 Jun 2026
Safety

Safety Alignment of LMs via Non-cooperative Games

DGX agent

arXiv:2512.20806v3 Announce Type: replace Abstract: Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rel

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

DGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

DGX agent

arXiv:2605.28830v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a c

model-releasesarxiv-cs-ai
29 May 2026
Safety

SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving

DGX agent

arXiv:2605.28583v1 Announce Type: cross Abstract: Ensuring both safety and efficiency in decision-making for autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement

safetyarxiv-cs-ai
28 May 2026
Safety

Curriculum Learning for Safety Alignment

DGX agent

arXiv:2605.26315v1 Announce Type: cross Abstract: Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibit

safetyarxiv-cs-ai
27 May 2026
Safety

AI-based Prediction of Independent Construction Safety Outcomes from Universal Attributes

DGX agent

arXiv:1908.05972v3 Announce Type: replace Abstract: This paper significantly improves on, and finishes to validate, an approach proposed in previous research in which safety outcomes were predicted fr

safetyarxiv-cs-lg
21 May 2026
Safety

PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment

DGX agent

arXiv:2605.21225v1 Announce Type: new Abstract: We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it f

safetyarxiv-cs-lg
21 May 2026
Safety

Distributional AGI Safety

DGX agent

arXiv:2512.16856v2 Announce Type: replace Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an e

safetyarxiv-cs-ai
20 May 2026
Safety

Reactive Robot-Centric Safety for Autonomous Navigation in Constrained and Dynamic Environments

DGX agent

arXiv:2605.15782v1 Announce Type: new Abstract: In this work, we address the problem of ensuring real-time safety in autonomous robot navigation, in spatially constrained dynamic environments, by util

safetyarxiv-cs-ro
18 May 2026
Model Releases

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

DGX agent

arXiv:2605.12386v1 Announce Type: new Abstract: Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are tempo

model-releasesarxiv-cs-ro
13 May 2026
Safety

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

DGX agent

arXiv:2601.18061v3 Announce Type: replace Abstract: Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI

safetyarxiv-cs-ai
12 May 2026
Safety

GuardAD: Safeguarding Autonomous Driving MLLMs via Markovian Safety Logic

DGX agent

arXiv:2605.10386v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse sa

safetyarxiv-cs-ai
12 May 2026
Safety

Why Does Agentic Safety Fail to Generalize Across Tasks?

DGX agent

arXiv:2605.06992v1 Announce Type: new Abstract: AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen

safetyarxiv-cs-lg
11 May 2026
Safety

Safety by Invariance, Liveness through Refinement: Heterogeneous Contract Framework for Co-Design of Layered Control

DGX agent

arXiv:2605.04222v1 Announce Type: cross Abstract: Real-world control systems must achieve long-horizon objectives (liveness) while respecting continuous-time safety constraints, a combination that mot

safetyarxiv-cs-ro
7 May 2026
Safety

Safety-critical Control Under Partial Observability: Reach-Avoid POMDP meets Belief Space Control

DGX agent

arXiv:2603.10572v2 Announce Type: replace Abstract: Partially Observable Markov Decision Processes (POMDPs) provide a principled framework for robot decision-making under uncertainty. Solving reach-av

safetyarxiv-cs-ro
6 May 2026
Safety

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

DGX agent

arXiv:2605.00227v1 Announce Type: new Abstract: There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely

safetyarxiv-cs-cl
4 May 2026
← Previous
1…45678…297
Next →