AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

DGX agent

arXiv:2605.29340v1 Announce Type: new Abstract: In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual an

safetyarxiv-cs-cl
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

DGX agent

arXiv:2605.29801v1 Announce Type: new Abstract: Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhi

model-releasesarxiv-cs-ai
29 May 2026
Safety

Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning

DGX agent

arXiv:2605.28552v1 Announce Type: new Abstract: As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle ty

safetyarxiv-cs-ai
28 May 2026
Model Releases

AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

DGX agent

arXiv:2605.26954v1 Announce Type: new Abstract: Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved.

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Position: AI Safety Requires Effective Controllability

DGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

model-releasesarxiv-cs-ai
27 May 2026
Safety

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

DGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

safetyarxiv-cs-ai
27 May 2026
Model Releases

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

DGX agent

arXiv:2605.24154v1 Announce Type: new Abstract: Current safety alignment of foundation models largely follows a one-size-fits-all paradigm, applying the same refusal policy across users and contexts.

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

DGX agent

arXiv:2601.21094v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safe

model-releasesarxiv-cs-ai
26 May 2026
Safety

CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety

DGX agent

arXiv:2605.21609v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating information seeking, advice, and emotionally sensit

safetyarxiv-cs-cl
22 May 2026
Safety

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

DGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

safetyarxiv-cs-cv
22 May 2026
Model Releases

Safety-Critical Control for Smoothed Implicit Contact Dynamics

DGX agent

arXiv:2605.21138v1 Announce Type: new Abstract: Smoothed implicit contact dynamics enables gradient-based planning and control for contact-rich tasks without predefined mode sequences. However, safety

model-releasesarxiv-cs-ro
21 May 2026
Model Releases

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

DGX agent

arXiv:2605.17971v1 Announce Type: cross Abstract: Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuri

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks

DGX agent

arXiv:2603.04459v3 Announce Type: replace-cross Abstract: The rapid expansion of research in LLM safety presents challenges in tracking advancements, making benchmarks important evaluation infrastruct

model-releasesarxiv-cs-ai
18 May 2026
Safety

parallelcbf: A composable safety-filter and auditability framework for tensor-parallel reinforcement learning

DGX agent

arXiv:2605.15509v1 Announce Type: new Abstract: While Isaac Lab provides massive parallel UAV simulation, OmniSafe and safe-control-gym provide constrained-RL benchmarks, and CBFKit provides control-b

safetyarxiv-cs-lg
18 May 2026
Safety

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

DGX agent

arXiv:2605.07250v1 Announce Type: cross Abstract: Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we i

safetyarxiv-cs-ai
11 May 2026
Model Releases

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

DGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

model-releasesarxiv-cs-ai
11 May 2026
Safety

Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content

DGX agent

arXiv:2605.04085v1 Announce Type: cross Abstract: Objectives: Large language models (LLMs) are increasingly used for clinical text summarization, yet structured methods to assess associated patient sa

safetyarxiv-cs-cl
7 May 2026
Safety

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments

DGX agent

arXiv:2508.04204v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have demonstrated impressive performance in reasoning-intensive tasks, but they remain vulnerable to harmful content g

safetyarxiv-cs-cl
7 May 2026
Safety

Improving Model Safety by Targeted Error Correction

DGX agent

arXiv:2605.02544v1 Announce Type: cross Abstract: The widespread adoption of machine learning in critical applications demands techniques to mitigate high-consequence errors. Our method utilizes a dua

safetyarxiv-cs-cv
5 May 2026
Model Releases

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

DGX agent

arXiv:2605.00689v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

DGX agent

arXiv:2604.26959v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into patient-facing healthcare systems offers significant potential to improve access to medical information.

model-releasesarxiv-cs-ai
1 May 2026
Safety

GAVEL: Towards Rule-Based Safety Through Activation Monitoring

DGX agent

arXiv:2601.19768v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be appare

safetyarxiv-cs-ai
1 May 2026
Safety

Test-Time Safety Alignment

DGX agent

arXiv:2604.26167v1 Announce Type: cross Abstract: Recent work has shown that a model's input word embeddings can serve as effective control variables for steering its behavior toward outputs that sati

safetyarxiv-cs-ai
30 Apr 2026
Safety

A Lightweight Explainable Guardrail for Prompt Safety

DGX agent

arXiv:2602.15853v2 Announce Type: replace-cross Abstract: We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly l

safetyarxiv-cs-ai
28 Apr 2026
Safety

Does Machine Unlearning Preserve Clinical Safety? A Risk Analysis for Medical Image Classification

DGX agent

arXiv:2604.23854v1 Announce Type: new Abstract: The application of Deep Learning in medical diagnosis must balance patient safety with compliance with data protection regulations. Machine Unlearning e

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models

DGX agent

arXiv:2511.08484v2 Announce Type: replace Abstract: We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilitie

model-releasesarxiv-cs-ai
28 Apr 2026
Safety

Time-Series Forecasting in Safety-Critical Environments: An EU-AI-Act-Compliant Open-Source Package / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein KI-VO-konformes Open-Source-Paket

DGX agent

arXiv:2604.23859v1 Announce Type: new Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical envir

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

DGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs

DGX agent

arXiv:2509.05367v4 Announce Type: replace-cross Abstract: Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classificati

safetyarxiv-cs-ai
17 Apr 2026
Safety

Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems

DGX agent

arXiv:2510.14133v2 Announce Type: replace Abstract: Agentic AI systems, which leverage multiple autonomous agents and large language models (LLMs), are increasingly used to address complex, multi-step

safetyarxiv-cs-ai
17 Apr 2026
Safety

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

DGX agent

arXiv:2604.12500v1 Announce Type: new Abstract: Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the condi

safetyarxiv-cs-lg
15 Apr 2026
Model Releases

Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking

DGX agent

arXiv:2604.10299v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) rely on attention-based retrieval of safety instructions to maintain alignment during generation. Existing attack

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

DGX agent

arXiv:2604.02022v2 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

DGX agent

arXiv:2604.06205v1 Announce Type: cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. Wh

safetyarxiv-cs-ai
10 Apr 2026
Safety

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

DGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

safetyarxiv-cs-ai
11 Aug 2026
Safety

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

DGX agent

arXiv:2608.08077v1 Announce Type: new Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

DGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs

DGX agent

arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, l

model-releasesarxiv-cs-cl
7 Aug 2026
Safety

Integrated Noise and Safety Management in UAM via A Unified Reinforcement Learning Framework

DGX agent

arXiv:2508.16440v2 Announce Type: replace-cross Abstract: Urban Air Mobility (UAM) envisions the widespread use of small aerial vehicles to transform transportation in dense urban environments. Howeve

safetyarxiv-cs-lg
6 Aug 2026
Model Releases

Item Response Theory for AI Safety

DGX agent

arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to tr

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

DGX agent

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-C

safetyarxiv-cs-cl
5 Aug 2026
Safety

MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection

DGX agent

arXiv:2608.00975v1 Announce Type: new Abstract: Large vision--language models (VLMs) can reason step by step about complex visual scenes, but this open-ended, autoregressive chain-of-thought (CoT) app

safetyarxiv-cs-cv
4 Aug 2026
Safety

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

DGX agent

arXiv:2607.26170v1 Announce Type: new Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting imag

safetyarxiv-cs-cv
30 Jul 2026
Safety

Safety-Aware Cascaded Inference for Crop Damage Assessment with Controlled Error Trade-offs

DGX agent

arXiv:2607.25468v1 Announce Type: new Abstract: In picture-based agricultural insurance for smallholder farmers, missed damage detections carry substantially higher cost than false alarms: a farmer wh

safetyarxiv-cs-cv
29 Jul 2026
Model Releases

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

DGX agent

arXiv:2607.22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authoriz

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Distributed Motion Planning with Safety Guarantees for Self-Reconfiguring Robotic Boats

DGX agent

arXiv:2607.20352v1 Announce Type: new Abstract: Aquatic self-reconfigurable robots must assemble into desired shapes while ensuring safe interactions among multiple agents. This paper proposes a hybri

safetyarxiv-cs-ro
23 Jul 2026
Model Releases

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

DGX agent

arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility a

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Adversarial Prompting Framework for AI Safety Assessment

DGX agent

arXiv:2607.13453v1 Announce Type: cross Abstract: Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use o

safetyarxiv-cs-ai
16 Jul 2026
← Previous
1…7891011…255
Next →