AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

Containment Verification: AI Safety Guarantees Independent of Alignment

DGX agent

arXiv:2605.09045v1 Announce Type: new Abstract: Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain

safetyarxiv-cs-ai
12 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

DGX agent

arXiv:2605.02900v1 Announce Type: cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, saf

safetyarxiv-cs-cv
6 May 2026
Safety

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection

DGX agent

arXiv:2509.00673v2 Announce Type: replace Abstract: We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety

safetyarxiv-cs-cl
5 May 2026
Safety

TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks

DGX agent

arXiv:2605.01761v1 Announce Type: new Abstract: Text-to-Video (T2V) models have demonstrated remarkable capability in generating temporally coherent videos from natural language prompts, yet they also

safetyarxiv-cs-cv
5 May 2026
Safety

A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety

DGX agent

arXiv:2505.17342v2 Announce Type: replace Abstract: Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and

safetyarxiv-cs-lg
30 Apr 2026
Safety

From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Model

DGX agent

arXiv:2604.26052v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes such as attack success rate, refusal rate, or harmful/not-harmful re

safetyarxiv-cs-cl
30 Apr 2026
Model Releases

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

DGX agent

arXiv:2601.18491v2 Announce Type: replace Abstract: The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current gua

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

Reasoning Structure Matters for Safety Alignment of Reasoning Models

DGX agent

arXiv:2604.18946v1 Announce Type: new Abstract: Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This

safetyarxiv-cs-ai
22 Apr 2026
Safety

Vision-Based Human Awareness Estimation for Enhanced Safety and Efficiency of AMRs in Industrial Warehouses

DGX agent

arXiv:2604.18627v1 Announce Type: new Abstract: Ensuring human safety is of paramount importance in warehouse environments that feature mixed traffic of human workers and autonomous mobile robots (AMR

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains

DGX agent

arXiv:2604.19001v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm

model-releasesarxiv-cs-cl
22 Apr 2026
Safety

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

DGX agent

arXiv:2604.16424v1 Announce Type: cross Abstract: State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

DGX agent

arXiv:2601.04043v2 Announce Type: replace Abstract: As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

DGX agent

arXiv:2509.18127v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, un

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies

DGX agent

arXiv:2604.09189v1 Announce Type: cross Abstract: LLMs internalize safety policies through RLHF, yet these policies are never formally specified and remain difficult to inspect. Existing benchmarks ev

safetyarxiv-cs-ai
13 Apr 2026
Safety

OpenKedge: Governing Agentic Mutation with Execution-Bound Safety and Evidence Chains

DGX agent

arXiv:2604.08601v1 Announce Type: new Abstract: The rise of autonomous AI agents exposes a fundamental flaw in API-centric architectures: probabilistic systems directly execute state mutations without

safetyarxiv-cs-ai
13 Apr 2026
Safety

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

DGX agent

arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis

safetyarxiv-cs-cl
12 Aug 2026
Safety

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

DGX agent

arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surroun

safetyarxiv-cs-ai
12 Aug 2026
Safety

Robust Safety Filtering for Input-Constrained Underactuated Linear Systems

DGX agent

arXiv:2608.10872v1 Announce Type: cross Abstract: We present a robust safety-filtering framework for input-constrained underactuated linear systems subject to unknown disturbances. A baseline H-infty

safetyarxiv-cs-ro
12 Aug 2026
Safety

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

DGX agent

arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep est

safetyarxiv-cs-cl
7 Aug 2026
Safety

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

DGX agent

arXiv:2608.04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research ha

safetyarxiv-cs-ai
6 Aug 2026
Safety

SCOPE: Field-of-View-Aware Path Planning in Unknown 3D Environments via Safety-Volume Certification

DGX agent

arXiv:2608.04420v1 Announce Type: new Abstract: Safe navigation with a body-mounted limited-field-of-view sensor requires the complete robot-inflated volume of an intended motion to be observed and ve

safetyarxiv-cs-ro
6 Aug 2026
Safety

ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories

DGX agent

arXiv:2608.03866v1 Announce Type: new Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework

safetyarxiv-cs-ai
5 Aug 2026
Safety

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

DGX agent

arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language

safetyarxiv-cs-cl
29 Jul 2026
Safety

Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety

DGX agent

arXiv:2607.24168v1 Announce Type: new Abstract: Road crashes remain among the gravest threats to public safety, and preventing them is a defining task of transportation systems worldwide. Much of that

safetyarxiv-cs-lg
28 Jul 2026
Safety

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

DGX agent

arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whe

safetyarxiv-cs-ai
28 Jul 2026
Safety

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

DGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

DGX agent

arXiv:2607.08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task

model-releasesarxiv-cs-ai
10 Jul 2026
Safety

Efficient Safety Alignment of Language Models via Latent Personality Traits

DGX agent

arXiv:2607.07918v1 Announce Type: cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Late

safetyarxiv-cs-ai
10 Jul 2026
Safety

VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

DGX agent

arXiv:2607.05180v1 Announce Type: cross Abstract: Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the ve

safetyarxiv-cs-cv
7 Jul 2026
Safety

Multi-modal Rail Crossing Safety Analysis

DGX agent

arXiv:2607.01365v1 Announce Type: cross Abstract: Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our abil

safetyarxiv-cs-ai
3 Jul 2026
Safety

FastBridge: Closing the Model-Based Realization Gap in Safety Filters on 3D Gaussian Splatting for Fast Quadrotor Flight

DGX agent

arXiv:2607.01200v1 Announce Type: new Abstract: Fast quadrotor flight requires safe obstacle avoidance under tight onboard compute limits. While 3D Gaussian Splatting (3DGS) provides a continuous, geo

safetyarxiv-cs-ro
2 Jul 2026
Model Releases

Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios

DGX agent

arXiv:2601.21173v2 Announce Type: replace-cross Abstract: With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in com

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

DGX agent

arXiv:2606.30256v1 Announce Type: new Abstract: Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides p

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

DGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Real-Time Safety Evaluation of Human Arm Operations Using a Wrist-Mounted IMU with PSM System

DGX agent

arXiv:2502.09241v2 Announce Type: replace Abstract: This paper presents a novel approach to real-time safety monitoring in human-robot collaborative manufacturing environments through a wrist-mounted

model-releasesarxiv-cs-ro
26 Jun 2026
Safety

Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications

DGX agent

arXiv:2606.23858v1 Announce Type: cross Abstract: A primary challenge in AI safety is the existence of adversarial examples -- slightly distorted inputs that cause a neural network (NN) to misclassify

safetyarxiv-cs-ai
24 Jun 2026
Model Releases

IndicGuard: A Multilingual Safety Guard Model and Dataset for Indic Languages

DGX agent

arXiv:2606.22841v1 Announce Type: cross Abstract: As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes, ensuring their safety and alignment with regional

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

DGX agent

arXiv:2606.22686v1 Announce Type: cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work, we investig

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

Benchmarking Large Language Models for Safety Data Extraction

DGX agent

arXiv:2606.11204v1 Announce Type: new Abstract: Accurate extraction of structured information from Safety Data Sheets (SDS) remains challenging in industrial safety due to heterogeneous document forma

model-releasesarxiv-cs-cl
11 Jun 2026
Safety

Schutzen: Evaluating LLM Safety in Bulgarian and German Contexts

DGX agent

arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disr

safetyarxiv-cs-cl
11 Jun 2026
Safety

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

DGX agent

arXiv:2606.09866v1 Announce Type: cross Abstract: Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods

safetyarxiv-cs-ai
10 Jun 2026
Safety

CARE: A Conformal Safety Layer for Medical Summarization

DGX agent

arXiv:2606.08969v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

DGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

DGX agent

arXiv:2606.08860v1 Announce Type: new Abstract: Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing from digital maps, creating safety risks f

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

DGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

model-releasesarxiv-cs-cv
5 Jun 2026
Safety

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

DGX agent

arXiv:2602.06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications,

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

DGX agent

arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models

DGX agent

arXiv:2606.00773v1 Announce Type: new Abstract: Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety-relevant tr

model-releasesarxiv-cs-ro
2 Jun 2026
← Previous
1…678910…255
Next →