AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

Online Safety Filter for Deformable Object Manipulation with Horizon Agnostic Neural Operators

DGX agent

arXiv:2605.01069v1 Announce Type: new Abstract: Safety critical control of robotic manipulation tasks involving deformable media such as fluids, cloth, and soft objects remains challenging because exi

safetyarxiv-cs-ro
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Grammar-Constrained Refinement of Safety Operational Rules Using Language in the Loop: What Could Go Wrong

DGX agent

arXiv:2604.23523v1 Announce Type: cross Abstract: Safety specifications in cyber-physical systems (CPS) capture the operational conditions the system must satisfy to operate safely within its intended

safetyarxiv-cs-ai
28 Apr 2026
Safety

SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs

DGX agent

arXiv:2604.20930v1 Announce Type: cross Abstract: Internal Safety Collapse (ISC) is a failure mode in which frontier LLMs, when executing legitimate professional tasks whose correct completion structu

safetyarxiv-cs-ai
24 Apr 2026
Safety

Continual Safety Alignment via Gradient-Based Sample Selection

DGX agent

arXiv:2604.17215v1 Announce Type: new Abstract: Large language models require continuous adaptation to new tasks while preserving safety alignment. However, fine-tuning on even benign data often compr

safetyarxiv-cs-lg
21 Apr 2026
Safety

On Safety Risks in Experience-Driven Self-Evolving Agents

DGX agent

arXiv:2604.16968v1 Announce Type: new Abstract: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self

safetyarxiv-cs-cl
21 Apr 2026
Safety

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints

DGX agent

arXiv:2604.12384v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) remains highly fragile during fine-tuning, where even benign adaptation can degrade pre-trained refusal

safetyarxiv-cs-ai
15 Apr 2026
Local Ai

The Illusion of Cross-Lingual Safety in Low-Resource Languages

DGX agent

arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. How

local-aiarxiv-cs-cl
12 Aug 2026
Safety

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

DGX agent

arXiv:2608.07535v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding an

safetyarxiv-cs-ai
11 Aug 2026
Safety

Safety Cost of Steering Vectors Is Separable and Reducible

DGX agent

arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally comprom

safetyarxiv-cs-cl
11 Aug 2026
Safety

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

DGX agent

arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, a

safetyarxiv-cs-lg
31 Jul 2026
Model Releases

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

DGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

model-releasesarxiv-cs-cl
28 Jul 2026
Safety

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

DGX agent

arXiv:2607.19913v1 Announce Type: new Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-orient

safetyarxiv-cs-ai
23 Jul 2026
Safety

Harnessing Textual Refusal Directions for Multimodal Safety

DGX agent

arXiv:2606.31876v1 Announce Type: new Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. B

safetyarxiv-cs-ai
1 Jul 2026
Safety

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion

DGX agent

arXiv:2606.25097v1 Announce Type: new Abstract: Speculative decoding accelerates inference by letting a draft model propose tokens for a target model to verify, raising a concrete safety question: at

safetyarxiv-cs-lg
25 Jun 2026
Safety

Towards a Bathroom-Centered Human-Building Digital Twin Framework for Indoor Safety Analysis

DGX agent

arXiv:2606.23292v2 Announce Type: replace-cross Abstract: Bathroom use is a critical safety challenge for older adults because wet surfaces, constrained layouts, limited support, and frequent posture

safetyarxiv-cs-ai
25 Jun 2026
Safety

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

DGX agent

arXiv:2606.25034v1 Announce Type: new Abstract: General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversaria

safetyarxiv-cs-cv
25 Jun 2026
Safety

Safe-SAGE: Social-Semantic Adaptive Guidance for Safe Engagement through Laplace-Modulated Poisson Safety Functions

DGX agent

arXiv:2603.05497v3 Announce Type: replace Abstract: Traditional safety-critical control methods, such as control barrier functions, suffer from semantic blindness, exhibiting the same behavior around

safetyarxiv-cs-ro
23 Jun 2026
Safety

Listening to the Workforce: Measuring Construction Worker Safety Attitudes from Social Media Discourse Using LLMs

DGX agent

arXiv:2606.04450v1 Announce Type: new Abstract: Worker safety attitudes are key determinants of whether protective practices are applied or bypassed on construction sites. Yet measuring them at scale

safetyarxiv-cs-cl
4 Jun 2026
Safety

Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections

DGX agent

arXiv:2605.27418v1 Announce Type: cross Abstract: The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transp

safetyarxiv-cs-ro
28 May 2026
Model Releases

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

DGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

model-releasesarxiv-cs-cl
28 May 2026
Safety

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

DGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

safetyarxiv-cs-ai
26 May 2026
Safety

Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays

DGX agent

arXiv:2605.23351v1 Announce Type: new Abstract: We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while

safetyarxiv-cs-lg
25 May 2026
Safety

Generating Realistic Safety-Critical Scenarios for Vehicle-Pedestrian Interactions

DGX agent

arXiv:2605.17229v1 Announce Type: new Abstract: Automated driving system deployment requires rigorous validation across safety-critical vehicle-pedestrian interactions, yet real-world datasets rarely

safetyarxiv-cs-ro
19 May 2026
Safety

Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents

DGX agent

arXiv:2605.17830v1 Announce Type: new Abstract: Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under

safetyarxiv-cs-ai
19 May 2026
Safety

SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation

DGX agent

arXiv:2601.18442v3 Announce Type: replace Abstract: Autonomous Vehicle (AV) requires rigorous testing in safety-critical scenarios for safety validation, yet its validation is hindered by the high cos

safetyarxiv-cs-ro
19 May 2026
Safety

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability

DGX agent

arXiv:2605.14246v1 Announce Type: cross Abstract: Many safety-critical control problems are modeled as risk-sensitive partially observable Markov decision processes, where the controller must make dec

safetyarxiv-cs-ai
15 May 2026
Safety

Selective Safety Steering via Value-Filtered Decoding

DGX agent

arXiv:2605.14746v1 Announce Type: new Abstract: While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of wor

safetyarxiv-cs-lg
15 May 2026
Safety

Internalizing Safety Understanding in Large Reasoning Models via Verification

DGX agent

arXiv:2605.08930v1 Announce Type: new Abstract: While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment para

safetyarxiv-cs-ai
12 May 2026
Safety

Shields to Guarantee Probabilistic Safety in MDPs

DGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

safetyarxiv-cs-ai
12 May 2026
Safety

A Closed-Form Dual-Barrier CBF Safety Filter for Holonomic Robots on Incrementally Built Occupancy Grid Maps

DGX agent

arXiv:2605.05182v1 Announce Type: new Abstract: We present a dual-barrier control barrier function (CBF) safety filter for real-time, safety-critical velocity control of holonomic robots operating in

safetyarxiv-cs-ro
7 May 2026
Safety

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

DGX agent

arXiv:2605.01147v1 Announce Type: new Abstract: As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety propertie

safetyarxiv-cs-ai
6 May 2026
Safety

Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics

DGX agent

arXiv:2604.26836v1 Announce Type: new Abstract: Predictive safety filters (PSFs) leverage model predictive control to enforce constraint satisfaction during deep reinforcement learning (RL) exploratio

safetyarxiv-cs-lg
30 Apr 2026
Safety

Unifying Runtime Monitoring Approaches for Safety-Critical Machine Learning: Application to Vision-Based Landing

DGX agent

arXiv:2604.26411v1 Announce Type: new Abstract: Runtime monitoring is essential to ensure the safety of ML applications in safety-critical domains. However, current research is fragmented, with indepe

safetyarxiv-cs-lg
30 Apr 2026
Safety

LLM-Augmented Traffic Signal Control with LSTM-Based Traffic State Prediction and Safety-Constrained Decision Support

DGX agent

arXiv:2604.23902v1 Announce Type: new Abstract: Traffic signal control is a critical task in intelligent transportation systems, yet conventional fixed-time and rule-based methods often struggle to ad

safetyarxiv-cs-ai
28 Apr 2026
Safety

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

DGX agent

arXiv:2604.24348v1 Announce Type: new Abstract: The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS

safetyarxiv-cs-cl
28 Apr 2026
Safety

TSAssistant: A Human-in-the-Loop Agentic Framework for Automated Target Safety Assessment

DGX agent

arXiv:2604.23938v1 Announce Type: new Abstract: Target Safety Assessment (TSA) requires systematic integration of heterogeneous evidence, including genetic, transcriptomic, target homology, pharmacolo

safetyarxiv-cs-cl
28 Apr 2026
Safety

Cat-DPO: Category-Adaptive Safety Alignment

DGX agent

arXiv:2604.17299v1 Announce Type: new Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusin

safetyarxiv-cs-cl
21 Apr 2026
Safety

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

DGX agent

arXiv:2604.17730v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging d

safetyarxiv-cs-cl
21 Apr 2026
Safety

Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility

DGX agent

arXiv:2604.15579v1 Announce Type: cross Abstract: AI agents that interact with their environments through tools enable powerful applications, but in high-stakes business settings, unintended actions c

safetyarxiv-cs-ai
20 Apr 2026
Safety

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

DGX agent

arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

DGX agent

arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

DGX agent

arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original,

safetyarxiv-cs-cl
31 Jul 2026
Safety

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

DGX agent

arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk

safetyarxiv-cs-lg
30 Jul 2026
Safety

On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

DGX agent

arXiv:2607.23365v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education.

safetyarxiv-cs-ai
28 Jul 2026
Safety

Learning Personalized Safety Interventions for Haptic Human-Robot Shared Control

DGX agent

arXiv:2607.19534v1 Announce Type: new Abstract: Haptic feedback provides an implicit channel for communicating safety intentions during human-robot shared control. Existing haptic guidance systems typ

safetyarxiv-cs-ro
23 Jul 2026
Model Releases

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

DGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

model-releasesarxiv-cs-ai
15 Jul 2026
Safety

PB-OEL: A Performance-Bounded Online Ensemble Learning Framework With Mixed Feedback for Real-Time Safety Assessment

DGX agent

arXiv:2503.15581v2 Announce Type: replace Abstract: Real-time safety assessment is critical for ensuring the reliable operation of complex dynamic systems. However, obtaining full safety labels in rea

safetyarxiv-cs-lg
9 Jul 2026
Safety

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

DGX agent

arXiv:2607.05407v1 Announce Type: cross Abstract: Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child

safetyarxiv-cs-ai
8 Jul 2026
← Previous
12345…255
Next →