AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
13 Apr 2026

SafeMind: A Risk-Aware Differentiable Control Framework for Adaptive and Safe Quadruped Locomotion

SafetyDGX agent

arXiv:2604.09474v1 Announce Type: cross Abstract: Learning-based quadruped controllers achieve impressive agility but typically lack formal safety guarantees under model uncertainty, perception noise,

10 Apr 2026

A Giant-Step Baby-Step Classifier For Scalable and Real-Time Anomaly Detection In Industrial Control Systems and Water Treatment Systems

SafetyDGX agent

arXiv:2504.20906v4 Announce Type: replace-cross Abstract: The continuous monitoring of the interactions between cyber-physical components of any industrial control system (ICS) is required to secure a

12 Aug 2026

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

11 Aug 2026

Graph-Guided Safe Diffuser: Topological Graph Guidance for Safe Diffusion Planning

SafetyDGX agent

arXiv:2608.09484v1 Announce Type: new Abstract: Many diffusion-based planners enforce safety through inference-time guidance, but such interleaved trajectory deformations often degrade kinematic feasi

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

SafetyDGX agent

arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Howev

5 Aug 2026

Forbidden Region Dynamic Active Constraints in Robot-Assisted Minimally Invasive Surgery

SafetyDGX agent

arXiv:2608.03010v1 Announce Type: new Abstract: In robot-assisted surgery, Forbidden Region Active Constraints (FRAC) represent a control strategy that helps maintain task safety by generating anisotr

31 Jul 2026

Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling

SafetyDGX agent

arXiv:2607.27511v1 Announce Type: new Abstract: Imitation learning has shown increasing promise for autonomous robotic surgery, yet safe deployment remains challenging due to the safety-critical natur

Uncertainty quantification for trustworthy deep learning: Methods and measures

SafetyDGX agent

arXiv:2607.28248v1 Announce Type: cross Abstract: The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures

30 Jul 2026

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

SafetyDGX agent

arXiv:2607.26574v1 Announce Type: cross Abstract: Safety classifiers ('guards') are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning:

Steering Instruction Hierarchies at Inference Time

SafetyDGX agent

arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override confl

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, man…

SafetyDGX agent

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, many forefront members of the AI-Safety community, in their fer

29 Jul 2026

Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail

SafetyDGX agent

arXiv:2607.25664v1 Announce Type: new Abstract: Machine learning demand forecasts optimize statistical accuracy yet leave excess operational volatility that inflates safety stock and amplifies the Bul

LLM Scheming Inversely Scales with Pretraining Language Coverage

SafetyDGX agent

arXiv:2607.24769v1 Announce Type: new Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has emp

28 Jul 2026

Constrained Reinforcement Learning Using Successor Representations

SafetyDGX agent

arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to int

Physical AI Governance: From Theory to Practice Across Life Cycle

SafetyDGX agent

arXiv:2607.22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact wit

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

SafetyDGX agent

arXiv:2607.24199v1 Announce Type: new Abstract: Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversi

24 Jul 2026

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and…

SafetyDGX agent

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety an

16 Jul 2026

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

Model ReleasesDGX agent

arXiv:2607.13348v1 Announce Type: new Abstract: Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Pr

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

SafetyDGX agent

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop

10 Jul 2026

Persona Cartography: Charting Language Model Personality Traits in Weight Space

SafetyDGX agent

arXiv:2607.07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decompo

8 Jul 2026

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

SafetyDGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

2 Jul 2026

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

SafetyDGX agent

arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image im

1 Jul 2026

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

SafetyDGX agent

arXiv:2504.12557v3 Announce Type: replace-cross Abstract: Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In m

30 Jun 2026

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

SafetyDGX agent

arXiv:2603.05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which int

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

SafetyDGX agent

arXiv:2606.29115v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoi

26 Jun 2026

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

SafetyDGX agent

arXiv:2606.26151v1 Announce Type: cross Abstract: While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventi

25 Jun 2026

Causality-Based Parametric Control Barrier Function for Safe Multi-Vehicle Interaction

SafetyDGX agent

arXiv:2606.25134v1 Announce Type: new Abstract: Safe control has been widely studied in various safety-critical applications, for instance, autonomous driving. In order to ensure the autonomous vehicl

From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

SafetyDGX agent

arXiv:2502.02205v4 Announce Type: replace Abstract: The application of deep learning for partial differential equation (PDE)-constrained control is gaining increasing attention. However, existing meth

24 Jun 2026

Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection

SafetyDGX agent

arXiv:2512.15503v3 Announce Type: replace-cross Abstract: Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle fo

23 Jun 2026

Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking

SafetyDGX agent

arXiv:2606.22278v1 Announce Type: cross Abstract: Ensuring safety of learning-enabled robotic manipulation across diverse embodiments and tasks still requires significant manual engineering. Existing

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics

SafetyDGX agent

arXiv:2606.22319v1 Announce Type: cross Abstract: Robotic ultrasound scanning in real clinical environments requires both high-level clinical workflow reasoning and low-level closed-loop execution. Ph

Toward Machine Risk Perception: Integrating Trust Calibration and Precursor-Based Risk Estimation for Humanoid

SafetyDGX agent

arXiv:2606.20748v1 Announce Type: new Abstract: Humanoid robots are emerging as co-workers in smart manufacturing, yet their dynamic, human-like movements introduce safety risks that differ fundamenta

10 Jun 2026

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

SafetyDGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

9 Jun 2026

Emergent alignment and the projectability of ethical personas

SafetyDGX agent

arXiv:2606.09475v1 Announce Type: new Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection'

Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human

SafetyDGX agent

arXiv:2606.08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approva

Safe, Fluent and Acceptable Motion Generation and Execution for Human--Robot Interaction in Manufacturing Environments

SafetyDGX agent

arXiv:2606.08741v1 Announce Type: new Abstract: Robots operating in human environments must not only ensure physical safety but also exhibit behaviors that are understandable, fluent, and acceptable t

6 Jun 2026

Conformal Risk-Averse Decision Making with Action Conditional Guarantee

SafetyDGX agent

arXiv:2606.05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety g

4 Jun 2026

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

SafetyDGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

2 Jun 2026

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

SafetyDGX agent

arXiv:2606.01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may

The Harsh Truth: Segment-Level Analysis of Harsh Driving Events in Milan Using Large-Scale Telematics, Street Networks, and Google Street View

SafetyDGX agent

arXiv:2606.00261v1 Announce Type: new Abstract: Police-reported crash statistics remain the standard input for urban road-safety assessment, but their incompleteness and reporting lag limit their usef

29 May 2026

Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models

SafetyDGX agent

arXiv:2604.17176v2 Announce Type: replace-cross Abstract: Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajecto

Robust and Efficient Guardrails with Latent Reasoning

Model ReleasesDGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

28 May 2026

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

SafetyDGX agent

arXiv:2605.27766v1 Announce Type: new Abstract: LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongsi

No Safe Dose: How Training Data Drives Unsafe Image Generation

SafetyDGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

27 May 2026

Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning

SafetyDGX agent

arXiv:2605.26452v1 Announce Type: cross Abstract: Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints duri

25 May 2026

CBANet: A Compact Attention-Based CNN-BiLSTM Network for Aggressive Driving Event Detection

SafetyDGX agent

arXiv:2605.23471v1 Announce Type: cross Abstract: Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promisin

MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems

SafetyDGX agent

arXiv:2602.04431v2 Announce Type: replace Abstract: LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fai

23 May 2026

Kernel-Based Safe Exploration in Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.22207v1 Announce Type: cross Abstract: Safety has been a major concern when deploying deep reinforcement learning algorithms in the real world. A promising direction that ensures that the l

21 May 2026

Automatically Learning Construction Injury Precursors from Text

SafetyDGX agent

arXiv:1907.11769v4 Announce Type: replace Abstract: In light of the increasing availability of digitally recorded safety reports in the construction industry, it is important to develop methods to exp

Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions

SafetyDGX agent

arXiv:2605.20566v1 Announce Type: new Abstract: Active perception in uncertain environments requires robots to navigate safely while acquiring informative observations to reduce map uncertainty. These

19 May 2026

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

SafetyDGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

Scenario Generation in Roundabouts with Adjustable Interaction Intensity

SafetyDGX agent

arXiv:2605.18026v1 Announce Type: new Abstract: Roundabouts, characterized by frequent merging and yielding interactions, remain a safety-critical corner case for the development and testing of intell

18 May 2026

PCASim: Promptable Closed-loop Adversarial Simulation for Urban Traffic Environment

SafetyDGX agent

arXiv:2605.15654v1 Announce Type: new Abstract: Real-world autonomous driving, particularly in urban environments with numerous corner cases, requires rigorous testing to ensure product safety and rob

SLIP & ETHICS: Graduated Intervention for AI Emotional Companions

SafetyDGX agent

arXiv:2605.15915v1 Announce Type: cross Abstract: AI emotional companions face a safety-rapport paradox: restrictive safeguards can damage supportive alliance, while permissive systems risk user harm.

15 May 2026

Training ML Models with Predictable Failures

SafetyDGX agent

arXiv:2605.15134v1 Announce Type: new Abstract: Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely l

14 May 2026

DisaBench: A Participatory Evaluation Framework for Disability Harms in Language Models

SafetyDGX agent

arXiv:2605.12702v1 Announce Type: new Abstract: General-purpose safety benchmarks for large language models do not adequately evaluate disability-related harms. We introduce DisaBench: a taxonomy of t

Revealing Interpretable Failure Modes of VLMs

SafetyDGX agent

arXiv:2605.12674v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to general

13 May 2026

Smart moves: Building resilient transportation systems with Google AI

SafetyDGX agent

What does transportation mean to you? For some, it’s making sure the train is on schedule so they can get to work on time. Maybe it’s making sure you have time connecting between flights. Maybe it’s a

12 May 2026

ASACK : Adaptive Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees

SafetyDGX agent

arXiv:2605.09659v1 Announce Type: new Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear operator acting on lifted observables, enabli

Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations

SafetyDGX agent

arXiv:2605.08754v1 Announce Type: new Abstract: Taxiway routing and on-surface conflict avoidance are coupled safety-critical decision problems in airport surface operations. Existing planning and opt

← Previous
1…1415161718…238
Next →