AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment

DGX agent

arXiv:2606.00369v1 Announce Type: cross Abstract: Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets rema

safetyarxiv-cs-lg
2 Jun 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

DGX agent

arXiv:2606.00611v1 Announce Type: new Abstract: Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moder

safetyarxiv-cs-ai
2 Jun 2026
Safety

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

DGX agent

arXiv:2605.31381v1 Announce Type: new Abstract: We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

DGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

model-releasesarxiv-cs-cv
28 May 2026
Safety

SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

DGX agent

arXiv:2605.28338v1 Announce Type: new Abstract: Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance

safetyarxiv-cs-ai
28 May 2026
Safety

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

DGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

safetyarxiv-cs-cl
28 May 2026
Safety

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

DGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

safetyarxiv-cs-ai
28 May 2026
Safety

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

DGX agent

arXiv:2605.26947v1 Announce Type: new Abstract: Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt datase

safetyarxiv-cs-cl
27 May 2026
Safety

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

DGX agent

arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align

safetyarxiv-cs-lg
26 May 2026
Safety

Towards Context-Invariant Safety Alignment for Large Language Models

DGX agent

arXiv:2605.20994v1 Announce Type: new Abstract: Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a stand

safetyarxiv-cs-cl
21 May 2026
Safety

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey

DGX agent

arXiv:2505.17352v2 Announce Type: replace Abstract: Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignmen

safetyarxiv-cs-cv
19 May 2026
Model Releases

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

DGX agent

arXiv:2605.16282v1 Announce Type: cross Abstract: The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a prolifera

model-releasesarxiv-cs-ai
19 May 2026
Safety

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

DGX agent

arXiv:2602.07892v2 Announce Type: replace-cross Abstract: Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility

safetyarxiv-cs-cl
13 May 2026
Safety

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models

DGX agent

arXiv:2605.08513v1 Announce Type: cross Abstract: Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expr

safetyarxiv-cs-ai
12 May 2026
Model Releases

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

DGX agent

arXiv:2503.06223v5 Announce Type: replace Abstract: Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is

model-releasesarxiv-cs-cv
11 May 2026
Safety

Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation

DGX agent

arXiv:2605.04366v1 Announce Type: cross Abstract: Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation off

safetyarxiv-cs-lg
7 May 2026
Safety

FORMULA: FORmation MPC with neUral barrier Learning for safety Assurance

DGX agent

arXiv:2604.04409v2 Announce Type: replace Abstract: Multi-robot systems (MRS) are essential for large-scale applications such as disaster response, material transport, and warehouse logistics, yet ens

safetyarxiv-cs-ro
6 May 2026
Model Releases

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

DGX agent

arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable t

model-releasesarxiv-cs-cl
5 May 2026
Safety

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

DGX agent

arXiv:2604.26577v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this

safetyarxiv-cs-ai
30 Apr 2026
Safety

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

DGX agent

arXiv:2603.02259v2 Announce Type: replace-cross Abstract: Multi-agent systems provide mature methodologies for role decomposition, coordination, and normative governance, capabilities that remain esse

safetyarxiv-cs-lg
30 Apr 2026
Safety

AI Safety Training Can be Clinically Harmful

DGX agent

arXiv:2604.23445v1 Announce Type: cross Abstract: Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigo

safetyarxiv-cs-ai
28 Apr 2026
Safety

Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems

DGX agent

arXiv:2604.22244v1 Announce Type: new Abstract: Ensuring safety for black-box hybrid dynamical systems presents significant challenges due to their instantaneous state jumps and unknown explicit nonli

safetyarxiv-cs-ro
27 Apr 2026
Safety

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem

DGX agent

arXiv:2506.17299v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their v

safetyarxiv-cs-ai
27 Apr 2026
Safety

Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression

DGX agent

arXiv:2505.13527v3 Announce Type: replace-cross Abstract: Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to j

safetyarxiv-cs-ai
24 Apr 2026
Safety

A Hough transform approach to safety-aware scalar field mapping using Gaussian Processes

DGX agent

arXiv:2604.20799v1 Announce Type: new Abstract: This paper presents a framework for mapping unknown scalar fields using a sensor-equipped autonomous robot operating in unsafe environments. The unsafe

safetyarxiv-cs-ro
23 Apr 2026
Safety

LLM-Guided Safety Agent for Edge Robotics with an ISO-Compliant Perception-Compute-Control Architecture

DGX agent

arXiv:2604.20193v1 Announce Type: new Abstract: Ensuring functional safety in human-robot interaction is challenging because AI perception is inherently probabilistic, whereas industrial standards req

safetyarxiv-cs-ro
23 Apr 2026
Safety

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

DGX agent

arXiv:2510.08240v2 Announce Type: replace Abstract: Harnessing the power of LLMs requires a delicate dance between being helpful and harmless. This creates a fundamental tension between two competing

safetyarxiv-cs-cl
22 Apr 2026
Safety

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

DGX agent

arXiv:2405.13068v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Cu

safetyarxiv-cs-lg
21 Apr 2026
Safety

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints

DGX agent

arXiv:2604.16916v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respo

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Hierarchical Reinforcement Learning with Runtime Safety Shielding for Power Grid Operation

DGX agent

arXiv:2604.14032v1 Announce Type: cross Abstract: Reinforcement learning has shown promise for automating power-grid operation tasks such as topology control and congestion management. However, its de

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

Goal-Conditioned Neural ODEs with Guaranteed Safety and Stability for Learning-Based All-Pairs Motion Planning

DGX agent

arXiv:2604.02821v2 Announce Type: replace Abstract: This paper presents a learning-based approach for all-pairs motion planning, where the initial and goal states are allowed to be arbitrary points in

safetyarxiv-cs-ro
15 Apr 2026
Model Releases

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

DGX agent

arXiv:2607.16943v2 Announce Type: replace Abstract: Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenari

model-releasesarxiv-cs-ro
12 Aug 2026
Model Releases

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

DGX agent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

DGX agent

arXiv:2608.08485v1 Announce Type: new Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inf

safetyarxiv-cs-ai
11 Aug 2026
Safety

Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems

DGX agent

arXiv:2608.06378v1 Announce Type: cross Abstract: Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly focus on drive

safetyarxiv-cs-ai
10 Aug 2026
Safety

JTA: Joint Testability Architecture for Scenario-Based Validation of Safety-Critical Software

DGX agent

arXiv:2608.05594v1 Announce Type: cross Abstract: Validation adequacy in safety-critical software depends on more than the system under test. Critical scenarios must be constructed under controlled co

safetyarxiv-cs-ro
7 Aug 2026
Safety

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

DGX agent

arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific

safetyarxiv-cs-cl
6 Aug 2026
Model Releases

S^3: Improving Agent Safety through Multi-Stage Defense

DGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

MROPE: A Multi-Robot Safe Cooperative Strategy via combined Predictive Safety Filters and Ellipse-based Constraint Compression

DGX agent

arXiv:2607.29203v1 Announce Type: new Abstract: Deploying drone swarms to track a dynamic target in cluttered environments presents severe computational and safety challenges. We propose MROPE, a hier

safetyarxiv-cs-ro
3 Aug 2026
Safety

Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning

DGX agent

arXiv:2607.27626v1 Announce Type: new Abstract: Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor's peak Age of Inform

safetyarxiv-cs-lg
31 Jul 2026
Safety

Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling

DGX agent

arXiv:2603.14841v3 Announce Type: replace-cross Abstract: Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offer l

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

DGX agent

arXiv:2607.15166v2 Announce Type: replace Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? W

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Matching Ranks Over Probability Yields Truly Deep Safety Alignment

DGX agent

arXiv:2512.05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their 'open' nature introduces more avenues fo

safetyarxiv-cs-ai
23 Jul 2026
Safety

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

DGX agent

arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature

safetyarxiv-cs-ai
16 Jul 2026
Safety

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

DGX agent

arXiv:2607.12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequentl

safetyarxiv-cs-ai
15 Jul 2026
Safety

Detecting Architectural Drift in Safety-Critical Firmware through Runtime Trace Analysis

DGX agent

arXiv:2607.03135v1 Announce Type: cross Abstract: Maintaining consistency between architectural design and runtime-observed behavior is challenging in long-lived safety-critical firmware. This paper p

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

DGX agent

arXiv:2606.20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Agentic Safety is an Epistemic Property, Not a Behavioral One

DGX agent

arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods

safetyarxiv-cs-ai
30 Jun 2026
← Previous
1…34567…255
Next →