AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

DGX agent

arXiv:2607.00402v1 Announce Type: cross Abstract: Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent metho

safetyarxiv-cs-ai
2 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

DGX agent

arXiv:2606.30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while th

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes

DGX agent

arXiv:2606.27210v1 Announce Type: new Abstract: We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS

safetyarxiv-cs-cl
26 Jun 2026
Safety

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning

DGX agent

arXiv:2606.05523v1 Announce Type: new Abstract: Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can byp

safetyarxiv-cs-cl
5 Jun 2026
Safety

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

DGX agent

arXiv:2606.04778v1 Announce Type: new Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent

safetyarxiv-cs-ai
4 Jun 2026
Safety

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

DGX agent

arXiv:2606.02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate t

safetyarxiv-cs-ai
2 Jun 2026
Safety

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization

DGX agent

arXiv:2510.09330v3 Announce Type: replace Abstract: Ensuring that large language models (LLMs) comply with safety requirements is a central challenge in AI deployment. Existing alignment approaches pr

safetyarxiv-cs-lg
2 Jun 2026
Safety

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

DGX agent

arXiv:2605.31073v1 Announce Type: new Abstract: Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do

safetyarxiv-cs-cl
1 Jun 2026
Safety

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection

DGX agent

arXiv:2605.28030v1 Announce Type: cross Abstract: Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversa

safetyarxiv-cs-ai
28 May 2026
Safety

BarrierSteer: LLM Safety via Learning Barrier Steering

DGX agent

arXiv:2602.20102v2 Announce Type: replace-cross Abstract: Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe co

safetyarxiv-cs-ai
25 May 2026
Safety

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

DGX agent

arXiv:2605.21834v1 Announce Type: new Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Con

safetyarxiv-cs-lg
23 May 2026
Safety

Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

DGX agent

arXiv:2605.14174v1 Announce Type: new Abstract: Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet

safetyarxiv-cs-ro
15 May 2026
Safety

Safety-Critical LiDAR-Inertial Odometry with On-Manifold Deterministic Protection Level

DGX agent

arXiv:2605.09383v1 Announce Type: new Abstract: In safety-critical scenarios, the protection level of the autonomous navigation system is crucial for enabling mobile robots to perform safe tasks. Howe

safetyarxiv-cs-ro
12 May 2026
Safety

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

DGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

safetyarxiv-cs-ai
12 May 2026
Model Releases

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

DGX agent

arXiv:2605.04572v1 Announce Type: cross Abstract: Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors l

model-releasesarxiv-cs-lg
7 May 2026
Safety

Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning

DGX agent

arXiv:2605.00667v1 Announce Type: new Abstract: Safety is a primary challenge in real-world reinforcement learning (RL). Formulating safety requirements as state-wise constraints has become a prominen

safetyarxiv-cs-lg
4 May 2026
Safety

Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under Uncertainty

DGX agent

arXiv:2604.27193v1 Announce Type: new Abstract: Automatic Emergency Braking (AEB) systems represent a safety-critical national interest, with the National Highway Traffic Safety Administration (NHTSA)

safetyarxiv-cs-ro
1 May 2026
Safety

Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems

DGX agent

arXiv:2604.27358v1 Announce Type: new Abstract: As large language model (LLM) agents are deployed in high-stakes environments, the question of how safely to delegate subtasks to specialized sub-agents

safetyarxiv-cs-ai
1 May 2026
Model Releases

Intent Laundering: AI Safety Datasets Are Not What They Seem

DGX agent

arXiv:2602.16729v3 Announce Type: replace-cross Abstract: We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isol

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging

DGX agent

arXiv:2503.17239v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) is a common practice to adapt generalist models to specialized domains. However, recent studies show

safetyarxiv-cs-ai
24 Apr 2026
Model Releases

Secure LLM Fine-Tuning via Safety-Aware Probing

DGX agent

arXiv:2505.16737v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable success across many applications, but their ability to generate harmful content raises s

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

DGX agent

arXiv:2604.19559v1 Announce Type: new Abstract: Construction workers are highly vulnerable to heat stress, yet tools that translate real-time physiological data into actionable safety intelligence rem

safetyarxiv-cs-ai
22 Apr 2026
Safety

Safety-Critical Contextual Control via Online Riemannian Optimization with World Models

DGX agent

arXiv:2604.19639v1 Announce Type: cross Abstract: Modern world models are becoming too complex to admit explicit dynamical descriptions. We study safety-critical contextual control, where a Planner mu

safetyarxiv-cs-ai
22 Apr 2026
Model Releases

Guardrails in Logit Space: Safety Token Regularization for LLM Alignment

DGX agent

arXiv:2604.17210v1 Announce Type: new Abstract: Fine-tuning well-aligned large language models (LLMs) on new domains often degrades their safety alignment, even when using benign datasets. Existing sa

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification

DGX agent

arXiv:2604.17391v1 Announce Type: cross Abstract: RISC-V is emerging as a viable platform for automotive-grade embedded computing, with recent ISO 26262 ASIL-D certifications demonstrating readiness f

safetyarxiv-cs-lg
21 Apr 2026
Safety

Why Agents Compromise Safety Under Pressure

DGX agent

arXiv:2603.14975v2 Announce Type: replace-cross Abstract: Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering

safetyarxiv-cs-cl
21 Apr 2026
Safety

Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety

DGX agent

arXiv:2604.13101v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into aviation safety decision-making represents a significant technological advancement, yet their sta

safetyarxiv-cs-ai
17 Apr 2026
Safety

Injecting Hallucinations in Autonomous Vehicles: A Component-Agnostic Safety Evaluation Framework

DGX agent

arXiv:2510.07749v2 Announce Type: replace Abstract: Perception failures in autonomous vehicles (AV) remain a major safety concern because they are the basis for many accidents. To study how these fail

safetyarxiv-cs-ro
12 Aug 2026
Model Releases

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

DGX agent

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

DGX agent

arXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech

model-releasesarxiv-cs-ai
10 Aug 2026
Local Ai

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

DGX agent

arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward wh

local-aiarxiv-cs-ai
5 Aug 2026
Safety

Toward Certified Functional Safety for Industrial Humanoid Robots: The Fail-Passive Gap and a Feasibility Study

DGX agent

arXiv:2608.02809v1 Announce Type: new Abstract: Industrial humanoid robots are constrained less by locomotion or manipulation capability than by the immaturity of functional safety certification for l

safetyarxiv-cs-ro
5 Aug 2026
Safety

Towards General Language-Conditioned Latent Safety Filters

DGX agent

arXiv:2608.00315v1 Announce Type: cross Abstract: Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified

safetyarxiv-cs-lg
4 Aug 2026
Model Releases

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

DGX agent

arXiv:2607.22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation,

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

DGX agent

arXiv:2510.03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastru

safetyarxiv-cs-ai
28 Jul 2026
Safety

CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation

DGX agent

arXiv:2607.04451v1 Announce Type: new Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating con

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

Safety Targeted Embedding Exploit via Refinement

DGX agent

arXiv:2607.01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-r

model-releasesarxiv-cs-ai
3 Jul 2026
Safety

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

DGX agent

arXiv:2606.29581v1 Announce Type: cross Abstract: Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluations usually

safetyarxiv-cs-ai
30 Jun 2026
Safety

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

DGX agent

arXiv:2510.16492v4 Announce Type: replace Abstract: As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While

safetyarxiv-cs-cl
29 Jun 2026
Model Releases

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

DGX agent

arXiv:2606.25442v1 Announce Type: new Abstract: Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. Ho

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

SHIELD: Safety on Humanoids via CBFs In Expectation on Learned Dynamics

DGX agent

arXiv:2505.11494v3 Announce Type: replace Abstract: Robot learning has produced remarkably effective ``black-box'' controllers for complex tasks such as dynamic locomotion on humanoids. Yet ensuring d

safetyarxiv-cs-ro
23 Jun 2026
Safety

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

DGX agent

arXiv:2606.06037v2 Announce Type: cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on m

safetyarxiv-cs-cl
10 Jun 2026
Safety

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

DGX agent

arXiv:2606.07678v1 Announce Type: cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data se

safetyarxiv-cs-ai
9 Jun 2026
Safety

Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering

DGX agent

arXiv:2606.07193v1 Announce Type: new Abstract: Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent duri

safetyarxiv-cs-ro
8 Jun 2026
Model Releases

Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring

DGX agent

arXiv:2603.25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs). However, unsafe events are rare in real-world CPS operations, creating an extreme

model-releasesarxiv-cs-lg
8 Jun 2026
Safety

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

DGX agent

arXiv:2606.03793v1 Announce Type: new Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial

safetyarxiv-cs-cl
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing

DGX agent

arXiv:2606.00686v1 Announce Type: new Abstract: The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmfu

safetyarxiv-cs-lg
2 Jun 2026
← Previous
123456…255
Next →