AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
9 Jul 2026

LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection

SafetyDGX agent

arXiv:2607.06592v1 Announce Type: cross Abstract: Object detectors have many applications in safety-critical systems, but they are known to be sensitive to worst-case perturbations such as adversarial

7 Jul 2026

Explainable Reinforcement Learning for Adaptive Traffic Signal Control

SafetyDGX agent

arXiv:2607.03703v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control. However, in safety-critical infrastructure like traf

Hope for the Best, Prepare for the Worst: Occlusion-Aware Contingency Planning for Autonomous Vehicles

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2607.03155v1 Announce Type: new Abstract: The deployment of autonomous vehicles in urban environments introduces significant safety challenges, particularly in scenarios with occlusions, where c

Resolving Primitive-Sharing Ambiguity in Long-Tailed TLS-Based Industrial MEP Point Cloud Segmentation via Spatial Context Constraints

SafetyDGX agent

arXiv:2601.19128v2 Announce Type: replace Abstract: In terrestrial laser scanning (TLS)-based mechanical, electrical, and plumbing (MEP) point cloud segmentation, safety-critical components such as re

6 Jul 2026

F1 in Britain: Automated software to blame for crushing expectations

SafetyDGX agent

Race control displayed a 'Safety Car in this lap' message during the 2026 British Grand Prix at Silverstone, raising expectations for a final lap restart that never materialized, as the message was er

3 Jul 2026

Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring

SafetyDGX agent

arXiv:2607.02121v1 Announce Type: cross Abstract: As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical. Gu

Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

SafetyDGX agent

arXiv:2607.01794v1 Announce Type: cross Abstract: With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection,

2 Jul 2026

Robust Operational Space Control with Conformal Disturbance Bounds for Safe Redundant Manipulation

SafetyDGX agent

arXiv:2607.00424v1 Announce Type: new Abstract: Redundant robotic manipulators operating in constrained and human-interactive environments require accurate task-space tracking together with rigorous s

1 Jul 2026

On Optimizing Multimodal Jailbreaks for Spoken Language Models

SafetyDGX agent

arXiv:2603.19127v2 Announce Type: replace Abstract: As Spoken Language Models (SLMs) integrate speech and text modalities, they inherit the safety vulnerabilities of their LLM backbone while introduci

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

SafetyDGX agent

arXiv:2602.03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrop

30 Jun 2026

A Gravitational Interpretation of Fine-Tuning Reversion

SafetyDGX agent

arXiv:2606.28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearne

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

SafetyDGX agent

arXiv:2602.21160v4 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a

Sparse Autoencoders are Capable LLM Jailbreak Mitigators

SafetyDGX agent

arXiv:2602.12418v2 Announce Type: replace-cross Abstract: Jailbreak attacks remain a persistent threat to large language model safety. We propose Context-Conditioned Delta Steering (CC-Delta), an SAE-

The Undecidability of Artificial General Intelligence (AGI) Alignment

SafetyDGX agent

arXiv:2606.28639v1 Announce Type: cross Abstract: This article establishes the foundational mathematical limits of Artificial General Intelligence (AGI) safety, proving that the core barrier is not th

26 Jun 2026

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

SafetyDGX agent

arXiv:2606.26205v1 Announce Type: new Abstract: Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records,

25 Jun 2026

A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

SafetyDGX agent

arXiv:2606.25380v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural context

ARTOO-DARTU: Studying AR-HRC With AR Obstruction Mitigation During a Warehouse Task

SafetyDGX agent

arXiv:2606.25202v1 Announce Type: cross Abstract: Human-robot collaboration (HRC) often requires robot intentions and internal states to be conveyed to users for task efficiency and safety. Recently,

Conformal Recovery-Deadline Certificates for Runtime Assurance of Adapting Controllers

SafetyDGX agent

arXiv:2606.25371v1 Announce Type: cross Abstract: Runtime assurance (RTA) protects a safety-critical system by switching from an advanced controller to a verified safe controller when a monitored cond

23 Jun 2026

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

SafetyDGX agent

arXiv:2606.20681v1 Announce Type: new Abstract: Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Co

HumanHalo -- Safe and Efficient 3D Navigation Among Humans via Minimally Conservative MPC

SafetyDGX agent

arXiv:2510.17525v3 Announce Type: replace Abstract: Safe and efficient robotic navigation among humans is essential for integrating robots into everyday environments. Most existing approaches focus on

10 Jun 2026

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Model ReleasesDGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

9 Jun 2026

Can Data Work be Reparative?

SafetyDGX agent

arXiv:2606.09408v1 Announce Type: cross Abstract: We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and b

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

SafetyDGX agent

arXiv:2606.09178v1 Announce Type: cross Abstract: Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target l

Diffuse AI Control on Fuzzy Tasks

SafetyDGX agent

arXiv:2606.08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfiel

8 Jun 2026

A Causal Probabilistic Framework for Perception-Informed Closed-Loop Simulation of Autonomous Driving

SafetyDGX agent

arXiv:2606.07186v1 Announce Type: new Abstract: Software-in-the-loop (SIL) simulation is a cornerstone for the validation of modern automotive safety functions. However, many current frameworks utiliz

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

SafetyDGX agent

arXiv:2606.07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

SafetyDGX agent

arXiv:2606.06875v1 Announce Type: new Abstract: Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the

6 Jun 2026

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

SafetyDGX agent

arXiv:2606.05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow

Where Should Knowledge Enter? A Layered Framework for Knowledge Infusion in Multimodal Iterative Generative Mo

SafetyDGX agent

arXiv:2606.06356v1 Announce Type: new Abstract: Multimodal generative models produce fluent outputs but remain unreliable when generation must respect structured, domain-specific, or safety-critical k

5 Jun 2026

Drishti AI-Event Guardian: An Intelligent Real-Time Crowd Monitoring and Emergency Response System for Mass Gathering Events

SafetyDGX agent

arXiv:2606.05185v1 Announce Type: cross Abstract: Mass gathering events are associated with critical safety incidents caused by insufficient crowd monitoring and inadequate emergency response coordina

4 Jun 2026

Instance-Level Post Hoc Uncertainty Quantification in Object Detection

SafetyDGX agent

arXiv:2606.04656v1 Announce Type: cross Abstract: Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for saf

Multi-Agent Next-Best-View Optimization for Risk-Averse Planning

SafetyDGX agent

arXiv:2606.04158v1 Announce Type: new Abstract: Multi-agent Next-Best-View (NBV) selection for safe path planning in uncertain and unknown environments requires informative, safety-aware, and efficien

3 Jun 2026

Easy-to-Use Shielding for Reinforcement Learning

SafetyDGX agent

arXiv:2606.03804v1 Announce Type: new Abstract: Safe exploration is a key challenge in Reinforcement Learning (RL) that aims to prevent agents from making harmful decisions while exploring their envir

Jailbreak Attack Initializations as Extractors of Compliance Directions

SafetyDGX agent

arXiv:2502.09755v4 Announce Type: replace-cross Abstract: Safety-aligned LLMs respond to prompts with either compliance or refusal, each corresponding to distinct directions in the model's activation

2 Jun 2026

Embedding Semantic Risk into Distance Fields and CBFs for Online Monocular Safe Control

SafetyDGX agent

arXiv:2606.01605v1 Announce Type: new Abstract: We propose an online monocular perception-to-control framework that embeds semantic risk into the distance field used by Control Barrier Function (CBF)-

Multi-Objective Reinforcement Learning for Tactical Decision Making for Trucks in Highway Traffic

SafetyDGX agent

arXiv:2601.18783v2 Announce Type: replace-cross Abstract: Balancing safety, efficiency, and operational costs in highway driving poses a challenging decision-making problem for heavy-duty vehicles. A

Predicted-Flow Control Barrier Functions for Real-Time Safe Optimal Control

Model ReleasesDGX agent

arXiv:2606.00297v1 Announce Type: cross Abstract: Control barrier functions (CBFs) provide real-time safety guarantees through pointwise conditions on the state. However, synthesizing a valid CBF is d

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

SafetyDGX agent

arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate

Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

SafetyDGX agent

arXiv:2606.00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential a

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer

SafetyDGX agent

arXiv:2602.02557v2 Announce Type: replace-cross Abstract: Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignmen

1 Jun 2026

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

SafetyDGX agent

arXiv:2602.07253v2 Announce Type: replace Abstract: Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existin

Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems

SafetyDGX agent

This newsletter issue discusses three key topics in AI development and safety: the challenges involved in overseeing and controlling advanced AI systems, empirical findings about how protein folding A

29 May 2026

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

SafetyDGX agent

arXiv:2512.11944v2 Announce Type: replace-cross Abstract: Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and int

28 May 2026

Evolving and Detecting Multi-Turn Deception using Geometric Signatures

SafetyDGX agent

arXiv:2605.27671v1 Announce Type: cross Abstract: Safety defenses for large language models (LLMs) are typically trained and evaluated on single-turn prompts, yet real attacks often unfold as indirect

High-Fidelity Industrial Crash Dynamics Prediction via Geometry-Aware Operator Learning with Memory-Efficient Low-Rank Attention

SafetyDGX agent

arXiv:2605.27758v1 Announce Type: cross Abstract: Automotive crashworthiness optimization remains a safety-critical challenge, requiring the management of large-scale nonlinear structural deformations

27 May 2026

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

SafetyDGX agent

arXiv:2605.27157v1 Announce Type: new Abstract: Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robust

26 May 2026

DBPnet: Damper Characteristics-Based Bayesian Physics-Informed Neural Network for Wheel Load Estimation

SafetyDGX agent

arXiv:2605.24860v1 Announce Type: cross Abstract: Advanced driver assistance systems (ADAS) play an important role in modern automotive intelligence, significantly enhancing vehicle safety and stabili

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

SafetyDGX agent

arXiv:2602.00491v2 Announce Type: replace Abstract: Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it r

25 May 2026

Empowering 9-1-1 Calltaking Training with Generative AI: Experiences and Lessons Learned

SafetyDGX agent

arXiv:2602.13241v2 Announce Type: replace-cross Abstract: Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a susta

General Hazard Detection

SafetyDGX agent

arXiv:2605.23304v1 Announce Type: new Abstract: Hazard, as an abstract concept, is typically defined through cognitive-level logical reasoning rather than concrete examples. In contrast, existing haza

22 May 2026

Safe and Steerable Geometric Motion Policies for Robotic Dexterous Manipulation

SafetyDGX agent

arXiv:2605.21811v1 Announce Type: new Abstract: Robotic dexterous manipulation requires continuously reconciling objectives and constraints defined on heterogeneous geometric spaces: a robot controlle

21 May 2026

Reinforcement Learning for Risk Adaptation via Differentiable CVaR Barrier Functions

SafetyDGX agent

arXiv:2605.21257v1 Announce Type: new Abstract: Planning through crowded environments under uncertain obstacle motions remains difficult, as stochastic interactions often induce overly conservative be

20 May 2026

Exploring and Developing a Pre-Model Safeguard with Draft Models

SafetyDGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails

SafetyDGX agent

arXiv:2510.13727v2 Announce Type: replace Abstract: Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-g

Generative Auto-Bidding with Unified Modeling and Exploration

SafetyDGX agent

arXiv:2605.19457v1 Announce Type: new Abstract: Automated bidding is central to modern digital advertising. Early rule-based methods lacked adaptability, while subsequent Reinforcement Learning approa

k-Inductive Neural Barrier Certificates for Unknown Nonlinear Dynamics

SafetyDGX agent

arXiv:2605.20108v1 Announce Type: cross Abstract: While conventional (k=1) discrete-time barrier certificate conditions impose strict safety constraints by requiring the function to be non-increasing

19 May 2026

AI Alignment Breaks at the Edge

SafetyDGX agent

arXiv:2602.20042v2 Announce Type: replace Abstract: General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. T

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

SafetyDGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

New Wide-Net-Casting Jailbreak Attacks Risk Large Models

SafetyDGX agent

arXiv:2605.17128v1 Announce Type: cross Abstract: Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexpl

15 May 2026

Bellman Value Decomposition for Task Logic in Safe Optimal Control

SafetyDGX agent

arXiv:2602.19532v2 Announce Type: replace Abstract: Real-world tasks involve nuanced combinations of goal and safety specifications. In high dimensions, the challenge is exacerbated: formal automata b

← Previous
1…1617181920…238
Next →