AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

DGX agent

arXiv:2608.08471v1 Announce Type: new Abstract: Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed

safetyarxiv-cs-ai
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics

DGX agent

arXiv:2608.05656v1 Announce Type: cross Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks fa

safetyarxiv-cs-ai
7 Aug 2026
Model Releases

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs

DGX agent

arXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

From Vessel Trajectories to Safety-Critical Encounter Scenarios: A Generative AI Framework for Autonomous Ship Digital Testing

DGX agent

arXiv:2603.28067v2 Announce Type: replace Abstract: Digital testing has emerged as a key paradigm for the development and verification of autonomous maritime navigation systems, yet the availability o

safetyarxiv-cs-lg
4 Aug 2026
Safety

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

DGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

safetyarxiv-cs-ai
3 Aug 2026
Safety

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

DGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

safetyarxiv-cs-cl
30 Jul 2026
Safety

Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture

DGX agent

arXiv:2607.24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can b

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

DGX agent

arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, re

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

DGX agent

arXiv:2607.21151v1 Announce Type: new Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuiti

safetyarxiv-cs-ai
24 Jul 2026
Safety

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

DGX agent

arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure

safetyarxiv-cs-ai
23 Jul 2026
Safety

PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner

DGX agent

arXiv:2603.10330v2 Announce Type: replace-cross Abstract: Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that lear

safetyarxiv-cs-ai
16 Jul 2026
Safety

Efficient Partitioning Method of Large-Scale Public Safety Spatio-Temporal Data based on Information Loss Constraints

DGX agent

arXiv:2306.12857v3 Announce Type: replace Abstract: The storage, management, and application of massive spatio-temporal data are widely used in practical scenarios, including public safety. However, d

safetyarxiv-cs-lg
10 Jul 2026
Safety

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

DGX agent

arXiv:2607.00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed a

safetyarxiv-cs-ai
2 Jul 2026
Safety

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

DGX agent

arXiv:2607.00334v1 Announce Type: new Abstract: Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

DGX agent

arXiv:2603.29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing ben

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

Agent Safety Is Action Alignment

DGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

DGX agent

arXiv:2606.28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains p

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

RAS: Measuring LLM Safety Through Refusal Alignment

DGX agent

arXiv:2606.25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

DGX agent

arXiv:2606.24245v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, th

safetyarxiv-cs-ai
24 Jun 2026
Model Releases

Quality Is Not a Safety Proxy Under Quantization

DGX agent

arXiv:2606.10154v1 Announce Type: new Abstract: Quantized checkpoints are often screened first with quality metrics and only later, if at all, with direct safety tests. This paper audits that shortcut

model-releasesarxiv-cs-lg
10 Jun 2026
Safety

Impedance MPC for Physical Human-Robot Interaction: Predictive Disturbance Rejection with Joint-Limit Safety

DGX agent

arXiv:2606.08281v1 Announce Type: new Abstract: Physical human-robot interaction (pHRI) demands simultaneous trajectory accuracy and compliant safety under unplanned contact. Classical impedance contr

safetyarxiv-cs-ro
9 Jun 2026
Safety

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation

DGX agent

arXiv:2606.08414v1 Announce Type: cross Abstract: Diffusion policies have achieved remarkable success in robotic manipulation, yet they often fail to satisfy strict physical constraints required for s

safetyarxiv-cs-ai
9 Jun 2026
Safety

Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety

DGX agent

arXiv:2606.05461v1 Announce Type: new Abstract: Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantifi

safetyarxiv-cs-ai
6 Jun 2026
Model Releases

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

DGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

DGX agent

arXiv:2605.30368v1 Announce Type: cross Abstract: Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SS

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

DGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

model-releasesarxiv-cs-cl
29 May 2026
Safety

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

DGX agent

arXiv:2605.28664v1 Announce Type: cross Abstract: Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are sca

safetyarxiv-cs-cl
28 May 2026
Model Releases

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

DGX agent

arXiv:2601.01627v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety

model-releasesarxiv-cs-ai
28 May 2026
Safety

Safety-Critical Adaptive Impedance Control via Nonsmooth Control Barrier Functions under State and Input Constraints

DGX agent

arXiv:2605.28367v1 Announce Type: new Abstract: Safe physical interaction is critical for deploying robotic manipulators in human-robot interaction and contact-rich tasks, where uncertainty, external

safetyarxiv-cs-ro
28 May 2026
Safety

CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment

DGX agent

arXiv:2603.07211v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes tra

safetyarxiv-cs-lg
27 May 2026
Safety

Constrained Meta Reinforcement Learning with Provable Test-Time Safety

DGX agent

arXiv:2601.21845v2 Announce Type: replace Abstract: Meta reinforcement learning (RL) allows agents to leverage experience across a distribution of tasks on which the agent can train at will, enabling

safetyarxiv-cs-lg
27 May 2026
Safety

Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection

DGX agent

arXiv:2605.27155v1 Announce Type: cross Abstract: Testing object detectors in safety-critical domains requires semantically meaningful probes beyond pixel-level corruptions. We present SemProbe, a too

safetyarxiv-cs-ai
27 May 2026
Safety

CarlaNCAP: A Framework for Quantifying the Safety of Vulnerable Road Users in Infrastructure-Assisted Collective Perception Using EuroNCAP Scenarios

DGX agent

arXiv:2512.11551v2 Announce Type: replace Abstract: The growing number of road users has significantly increased the risk of accidents in recent years. Vulnerable Road Users (VRUs) are particularly at

safetyarxiv-cs-ro
25 May 2026
Local Ai

Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries

DGX agent

arXiv:2605.21712v1 Announce Type: new Abstract: Transportation safety analysis requires integrating crash records, roadway attributes, and geospatial data through GIS-based workflows, but access remai

local-aiarxiv-cs-cl
22 May 2026
Safety

Anomaly-Informed Confidence Calibration for Vision-Based Safety Prediction

DGX agent

arXiv:2605.21109v1 Announce Type: new Abstract: Reliable confidence estimates are important for safely deploying vision-based controllers in autonomous racing, where safety predictions must be derived

safetyarxiv-cs-ro
21 May 2026
Model Releases

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry

DGX agent

arXiv:2605.20241v1 Announce Type: cross Abstract: Prompt-level safety probes for large language models use hidden-state representations to separate safe from unsafe prompts, but strong average detecti

model-releasesarxiv-cs-cl
21 May 2026
Safety

Time-To-Reach Separation and Safety Filtering for Safe, Fair, and Efficient Multi-Agent Coordination

DGX agent

arXiv:2605.20625v1 Announce Type: cross Abstract: Advanced Air Mobility (AAM) operations are expected to significantly increase aerial traffic in urban airspace, requiring autonomous traffic managemen

safetyarxiv-cs-ro
21 May 2026
Safety

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

DGX agent

arXiv:2605.17413v1 Announce Type: cross Abstract: Safety-aligned language models often refuse cybersecurity requests whose wording resembles misuse, even when the task is authorized and defensive. Thi

safetyarxiv-cs-ai
19 May 2026
Safety

LLM-Safety Evaluations Lack Robustness

DGX agent

arXiv:2503.02574v2 Announce Type: replace-cross Abstract: In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of n

safetyarxiv-cs-ai
19 May 2026
Safety

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

DGX agent

arXiv:2603.18178v2 Announce Type: replace-cross Abstract: The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-co

safetyarxiv-cs-ai
19 May 2026
Model Releases

IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia

DGX agent

arXiv:2603.17915v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are deployed in multilingual settings, their safety behavior in culturally diverse, low-resource languages rem

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

DGX agent

arXiv:2509.26100v2 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, exis

model-releasesarxiv-cs-ai
15 May 2026
Safety

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

DGX agent

arXiv:2605.15164v1 Announce Type: cross Abstract: This position paper argues that behavioural assurance, even when carefully designed, is being asked to carry safety claims it cannot verify. AI govern

safetyarxiv-cs-ai
15 May 2026
Model Releases

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

DGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

model-releasesarxiv-cs-ai
15 May 2026
Safety

Before the Last Token: Diagnosing Final-Token Safety Probe Failures

DGX agent

arXiv:2605.12726v1 Announce Type: new Abstract: Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompts can contain probe-visible unsafe evidence distribute

safetyarxiv-cs-lg
14 May 2026
Safety

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

DGX agent

arXiv:2602.07668v2 Announce Type: replace Abstract: The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state

safetyarxiv-cs-cv
13 May 2026
Safety

Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

DGX agent

arXiv:2605.11769v1 Announce Type: new Abstract: Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. Whi

safetyarxiv-cs-cl
13 May 2026
← Previous
1…56789…255
Next →