AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
9 Jun 2026

Code Is More Than Text: Uncertainty Estimation for Code Generation

SafetyDGX agent

arXiv:2606.09577v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as code generators, where silently wrong programs pose real safety and reliability risks. Relia

Distant Object Localisation from Noisy Image Segmentation Sequences

SafetyDGX agent

arXiv:2509.20906v3 Announce Type: replace Abstract: 3D object localisation based on a sequence of camera measurements is essential for safety-critical surveillance tasks, such as drone-based wildfire

Human-Centered Benchmarking of Driver Monitoring Models

SafetyDGX agent

arXiv:2606.08123v1 Announce Type: cross Abstract: Vision-based driver monitoring systems are increasingly deployed in safety-critical intelligent transportation settings, yet they are almost always co

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Hyperspectral Smoke Segmentation via Mixture of Prototypes

SafetyDGX agent

arXiv:2602.10858v2 Announce Type: replace Abstract: Smoke segmentation is critical for wildfire management and industrial safety applications. Traditional visible-light-based methods face limitations

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences

SafetyDGX agent

arXiv:2606.07629v1 Announce Type: cross Abstract: Current approaches to aligning large language models (LLMs) aggregate diverse human preferences into a single reward signal, effectively optimizing fo

LUNA-AD: Lightweight Uncertainty-Aware Language Model with Lifelong Learning for Autonomous Driving

SafetyDGX agent

arXiv:2606.08470v1 Announce Type: new Abstract: While large language models (LLMs) offer promising reasoning capabilities, their integration into safety-critical driving systems is hindered by limited

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07706v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated strong performance across multimodal tasks, yet their safety robustness remains an open challenge. Whi

On-the-fly hand-eye calibration for the da Vinci surgical robot

SafetyDGX agent

arXiv:2601.14871v2 Announce Type: replace Abstract: In Robot-Assisted Minimally Invasive Surgery (RMIS), accurate tool localization is crucial to ensure patient safety and successful task execution. H

QDS-SNN: Energy-efficient Quantum Deeply-Supervised Spiking Neural Network Algorithm for Traffic Sign Recognition

SafetyDGX agent

arXiv:2606.07657v1 Announce Type: cross Abstract: Traffic sign recognition is crucial for intelligent transportation and autonomous driving, as it can improve driving efficiency and ensure road safety

Scaling Neural Network Verification with Tensor Parallelism and Fully Sharded Data Parallelism

SafetyDGX agent

arXiv:2606.09377v1 Announce Type: cross Abstract: Formal neural network verification -- proving that a network satisfies safety properties for all inputs in a specified domain -- is bounded in practic

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

SafetyDGX agent

arXiv:2601.04266v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex m

The Confidence Trap: Calibration Attacks for Graph Neural Networks

SafetyDGX agent

arXiv:2606.08467v1 Announce Type: cross Abstract: While confidence calibration is essential for trustworthy decision-making in safety-critical applications, the robustness of calibrated GNNs to advers

Vessel Traffic Flow Prediction on Sparse Data via Spatio-Temporal Graph Neural Networks with a Learnable Tweedie Head

SafetyDGX agent

arXiv:2606.07694v1 Announce Type: new Abstract: Accurate vessel traffic flow prediction is crucial for smart port operations and navigational safety. However, maritime traffic flow data are often high

8 Jun 2026

Built to benefit everyone: our plan

SafetyDGX agent

OpenAI outlines its mission and strategic approach to developing artificial intelligence technologies designed to benefit humanity broadly. The plan likely covers OpenAI's commitments to safety resear

Latent-space Attacks for Refusal Evasion in Language Models

SafetyDGX agent

arXiv:2605.21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representat

Multi-Scale Feature Attention Network for Polymer Classification using THz Dual-Comb Spectroscopy

SafetyDGX agent

arXiv:2606.06554v1 Announce Type: cross Abstract: Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectroscopic tech

Residual-Controlled Multiplier Learning for Stochastic Constrained Decision-Making

SafetyDGX agent

arXiv:2606.07088v1 Announce Type: new Abstract: Stochastic constrained decision-making requires optimizing performance objectives while enforcing statistical requirements such as safety or fairness. H

Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach

SafetyDGX agent

arXiv:2510.09041v3 Announce Type: replace-cross Abstract: Deep reinforcement learning (DRL) has demonstrated remarkable success in developing autonomous driving policies. However, its vulnerability to

6 Jun 2026

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

SafetyDGX agent

arXiv:2606.06223v1 Announce Type: new Abstract: Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal mode

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems

SafetyDGX agent

arXiv:2606.06114v1 Announce Type: new Abstract: Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degrada

5 Jun 2026

Real-Time Threat Detection from Surveillance Cameras using Machine Learning

SafetyDGX agent

arXiv:2606.05708v1 Announce Type: new Abstract: Ensuring public safety in densely populated urban environments remains a critical challenge, necessitating the deployment of intelligent and automated v

4 Jun 2026

Expert-Aware Refusal Steering

SafetyDGX agent

arXiv:2606.04160v1 Announce Type: new Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed r

Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms

SafetyDGX agent

arXiv:2606.04767v1 Announce Type: cross Abstract: The robustness of deep neural networks is crucial for safety-critical deployments, yet existing evaluation methods are often attack-dependent and lack

Testing Neural Networks via Bayesian-Guided Exploration of Decision Landscapes

SafetyDGX agent

arXiv:2606.04314v1 Announce Type: new Abstract: As neural networks are increasingly deployed in safety-critical domains, testing is essential to evaluate and improve their reliability. Existing testin

3 Jun 2026

Assessing Region-Level EEG Contributions to Cognitive Workload Prediction

SafetyDGX agent

arXiv:2606.02598v1 Announce Type: new Abstract: Accurate and generalizable estimation of cognitive workload from electroencephalography (EEG) is critical for human-centered and safety-critical systems

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

SafetyDGX agent

arXiv:2606.02640v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to i

Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat Intelligence

SafetyDGX agent

arXiv:2606.02967v1 Announce Type: cross Abstract: The space industry is quietly building toward something nobody has fully reckoned with: orbital data centers running thousands of autonomous AI worklo

LAP: An Agent-to-Instrument Protocol for Autonomous Science

SafetyDGX agent

arXiv:2606.03755v1 Announce Type: new Abstract: Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execu

Learning Power Flow with Confidence: A Probabilistic Guarantee Framework for Voltage Risk

SafetyDGX agent

arXiv:2308.07867v4 Announce Type: replace-cross Abstract: The absence of formal performance guarantees in machine learning (ML) has limited its adoption for safety-critical power system applications,

SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory

SafetyDGX agent

arXiv:2511.16275v4 Announce Type: replace-cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language models (LLMs) in safety-critical scenarios, as it enables t

2 Jun 2026

Adversarial Attacks on Robot Localization Systems via Deep Feature Perturbation

SafetyDGX agent

arXiv:2606.01892v1 Announce Type: new Abstract: Robot localization systems are critical for autonomous navigation and safety. Adversarial perturbations can mislead these systems, resulting in mislocal

CEAR: Certified Ensemble Adversarial Robustness in DNNs

SafetyDGX agent

arXiv:2606.01437v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) are highly susceptible to adversarial perturbations, leading to extensive research on robustness for safety-critical appli

Constrained Whole-Body Tracking for Humanoid Robots

SafetyDGX agent

arXiv:2606.00374v1 Announce Type: new Abstract: Recent advances in reinforcement learning (RL) have demonstrated impressive whole-body agility for humanoid robots, yet ensuring safety and satisfying c

DriveAnchor: Progressive Anchor-based Flow Learning for Autonomous Driving Planning

SafetyDGX agent

arXiv:2606.00519v1 Announce Type: new Abstract: We present DriveAnchor, a three-stage framework for autonomous driving planning that achieves behavioral diversity, controllability, and safety in a com

From Cues to Horizons: Dynamic Risk Horizon Profiling for Trajectory Prediction

SafetyDGX agent

arXiv:2606.00857v1 Announce Type: cross Abstract: Accurate and reliable vehicle trajectory prediction is essential for safe autonomous driving. Recent studies have incorporated safety risk into trajec

IstGPT: LLM-based Anomaly Detection for Spatial-Temporal Graph in Industrial Systems

SafetyDGX agent

arXiv:2606.01691v1 Announce Type: cross Abstract: Industrial Internet systems face increasing threats from sophisticated industrial control system (ICS) attacks, resulting in critical safety incidents

LFA: Layer Feature Attention for Run-Time Introspection of 2D Object Detectors in Automated Driving

SafetyDGX agent

arXiv:2606.00372v1 Announce Type: new Abstract: Reliable object detection is critical for automated driving, yet even state-of-the-art detectors inevitably make errors that can compromise safety. Intr

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling

Model ReleasesDGX agent

arXiv:2505.17659v4 Announce Type: replace-cross Abstract: Safe and feasible trajectory planning is critical for real-world autonomous driving systems. However, existing learning-based planners rely he

SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models

SafetyDGX agent

arXiv:2601.14323v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed in safety-critical robotic applications, yet their security vulnerabilities rema

Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection

SafetyDGX agent

arXiv:2606.01896v1 Announce Type: cross Abstract: Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expensive, or bia

1 Jun 2026

Simulation of collision avoidance behavior in crowd movement by data-driven approach

SafetyDGX agent

arXiv:2605.31210v1 Announce Type: cross Abstract: Crowd movement simulation is essential for pedestrian safety management and facility layout optimization. Data-driven models enhance trajectory predic

Unsupervised Defect Detection for Surgical Instruments

SafetyDGX agent

arXiv:2509.21561v2 Announce Type: replace Abstract: Ensuring the safety of surgical instruments requires reliable detection of visual defects. However, manual inspection is prone to error, and existin

29 May 2026

DefSynUS: Real-time Patient-specific Intrahepatic Vessel Identification via Deformation-Aware CT-US Domain Adaptation

SafetyDGX agent

arXiv:2605.29570v1 Announce Type: new Abstract: Purpose: Laparoscopic ultrasound (LUS) enhances the safety of liver surgery by visualizing intrahepatic vessels in real-time. Still, vessel identificati

Low-Magnification SEM May Suffice: Interpretable Deep Learning for Multi-Scale Fracture-Cause Classification in Zirconia-Toughened Alumina

SafetyDGX agent

arXiv:2605.29798v1 Announce Type: new Abstract: Reliable identification of fracture origins in alumina matrix composite hip and knee implants is critical for quality assurance and patient safety, yet

Masked Diffusion Modeling for Anomaly Detection

SafetyDGX agent

arXiv:2605.30046v1 Announce Type: cross Abstract: Anomaly detection aims to identify samples that deviate from the nominal data distribution and is central to many safety-critical applications. Howeve

Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving

SafetyDGX agent

arXiv:2605.29138v1 Announce Type: cross Abstract: Latency-accuracy tradeoffs are fundamental in real-time applications of deep neural networks (DNNs) for cyber-physical systems. In autonomous driving,

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

SafetyDGX agent

arXiv:2605.29146v1 Announce Type: cross Abstract: Medication recommendation predicts medications for patient visits, but existing methods still face two key challenges. At the model level, traditional

SigmaMedStat: Temporal Signal Modeling for ICU False Alarm Reduction

SafetyDGX agent

arXiv:2605.29236v1 Announce Type: new Abstract: Alarm fatigue in intensive care units (ICUs) is a well documented patient safety crisis. Clinical monitors generate 350 or more alarms per patient per d

V2XCrafter: Learning to Generate Driving Scene Across Agents

SafetyDGX agent

arXiv:2605.29471v1 Announce Type: new Abstract: Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet

28 May 2026

AOE: Exhaustive Out-of-Distribution Detection via Recalibrating Outlier Labels

SafetyDGX agent

arXiv:2605.28021v1 Announce Type: new Abstract: Out-of-distribution (OOD) detection is essential for deploying machine learning models in open-world and safety-critical scenarios, where test inputs ma

COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving

SafetyDGX agent

arXiv:2604.00402v2 Announce Type: replace-cross Abstract: Developing robust models to accurately predict the trajectories of surrounding agents is fundamental to autonomous driving safety. However, mo

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

SafetyDGX agent

arXiv:2605.27823v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulti

OpenAI’s Frontier Governance Framework

SafetyDGX agent

OpenAI's Frontier Governance Framework outlines the organization's approach to managing risks associated with advanced AI systems, including safety, security, and responsible deployment practices. The

Provably Guaranteed Polytopic Uncertainty Quantification for SLAM

SafetyDGX agent

arXiv:2605.28172v1 Announce Type: new Abstract: In safety-critical robotics applications, guaranteed and practical uncertainty quantification (UQ) in perception is vital. Many existing works either of

Voluntary Collusion with Secret Tools in Competing LLM Agents

SafetyDGX agent

arXiv:2605.27593v1 Announce Type: new Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collus

27 May 2026

CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intelligence

SafetyDGX agent

arXiv:2605.26524v1 Announce Type: cross Abstract: Maritime intelligent transportation systems (MITS) are essential for ensuring navigation safety and efficiency in busy waterways. However, accurate ve

26 May 2026

Certified Robustness from Approximate Gaussian Mixture Structures in Pretrained Latent Spaces

SafetyDGX agent

arXiv:2605.25352v1 Announce Type: cross Abstract: Deep learning models are vulnerable to adversarial perturbations, raising important concerns for safety-critical deployment. Empirical defenses can ac

Measuring the Depth of LLM Unlearning via Activation Patching

SafetyDGX agent

arXiv:2605.24614v1 Announce Type: cross Abstract: Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target kn

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection

Model ReleasesDGX agent

arXiv:2605.24834v1 Announce Type: cross Abstract: Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversar

Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation

SafetyDGX agent

arXiv:2605.24535v1 Announce Type: cross Abstract: Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation intervention

← Previous
1…2122232425…240
Next →