AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4

DGX agent

This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat

safetyimport-ai
20 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility

DGX agent

arXiv:2604.15579v1 Announce Type: cross Abstract: AI agents that interact with their environments through tools enable powerful applications, but in high-stakes business settings, unintended actions c

safetyarxiv-cs-ai
20 Apr 2026
Safety

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

DGX agent

arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

DGX agent

arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

DGX agent

arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original,

safetyarxiv-cs-cl
31 Jul 2026
Safety

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

DGX agent

arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk

safetyarxiv-cs-lg
30 Jul 2026
Safety

On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

DGX agent

arXiv:2607.23365v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education.

safetyarxiv-cs-ai
28 Jul 2026
Safety

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot…

DGX agent

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot of it is about people that just really, even if it's slight

safetyclem-delangue--x
28 Jul 2026
Safety

Learning Personalized Safety Interventions for Haptic Human-Robot Shared Control

DGX agent

arXiv:2607.19534v1 Announce Type: new Abstract: Haptic feedback provides an implicit channel for communicating safety intentions during human-robot shared control. Existing haptic guidance systems typ

safetyarxiv-cs-ro
23 Jul 2026
Model Releases

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

DGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

model-releasesarxiv-cs-ai
15 Jul 2026
Safety

PB-OEL: A Performance-Bounded Online Ensemble Learning Framework With Mixed Feedback for Real-Time Safety Assessment

DGX agent

arXiv:2503.15581v2 Announce Type: replace Abstract: Real-time safety assessment is critical for ensuring the reliable operation of complex dynamic systems. However, obtaining full safety labels in rea

safetyarxiv-cs-lg
9 Jul 2026
Safety

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

DGX agent

arXiv:2607.05407v1 Announce Type: cross Abstract: Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child

safetyarxiv-cs-ai
8 Jul 2026
Safety

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

DGX agent

arXiv:2607.00402v1 Announce Type: cross Abstract: Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent metho

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

DGX agent

arXiv:2606.30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while th

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes

DGX agent

arXiv:2606.27210v1 Announce Type: new Abstract: We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS

safetyarxiv-cs-cl
26 Jun 2026
Safety

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning

DGX agent

arXiv:2606.05523v1 Announce Type: new Abstract: Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can byp

safetyarxiv-cs-cl
5 Jun 2026
Safety

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

DGX agent

arXiv:2606.04778v1 Announce Type: new Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent

safetyarxiv-cs-ai
4 Jun 2026
Safety

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

DGX agent

arXiv:2606.02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate t

safetyarxiv-cs-ai
2 Jun 2026
Safety

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization

DGX agent

arXiv:2510.09330v3 Announce Type: replace Abstract: Ensuring that large language models (LLMs) comply with safety requirements is a central challenge in AI deployment. Existing alignment approaches pr

safetyarxiv-cs-lg
2 Jun 2026
Safety

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

DGX agent

arXiv:2605.31073v1 Announce Type: new Abstract: Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do

safetyarxiv-cs-cl
1 Jun 2026
Safety

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection

DGX agent

arXiv:2605.28030v1 Announce Type: cross Abstract: Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversa

safetyarxiv-cs-ai
28 May 2026
Safety

BarrierSteer: LLM Safety via Learning Barrier Steering

DGX agent

arXiv:2602.20102v2 Announce Type: replace-cross Abstract: Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe co

safetyarxiv-cs-ai
25 May 2026
Safety

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

DGX agent

arXiv:2605.21834v1 Announce Type: new Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Con

safetyarxiv-cs-lg
23 May 2026
Safety

Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

DGX agent

arXiv:2605.14174v1 Announce Type: new Abstract: Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet

safetyarxiv-cs-ro
15 May 2026
Safety

Safety-Critical LiDAR-Inertial Odometry with On-Manifold Deterministic Protection Level

DGX agent

arXiv:2605.09383v1 Announce Type: new Abstract: In safety-critical scenarios, the protection level of the autonomous navigation system is crucial for enabling mobile robots to perform safe tasks. Howe

safetyarxiv-cs-ro
12 May 2026
Safety

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

DGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

safetyarxiv-cs-ai
12 May 2026
Model Releases

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

DGX agent

arXiv:2605.04572v1 Announce Type: cross Abstract: Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors l

model-releasesarxiv-cs-lg
7 May 2026
Safety

Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning

DGX agent

arXiv:2605.00667v1 Announce Type: new Abstract: Safety is a primary challenge in real-world reinforcement learning (RL). Formulating safety requirements as state-wise constraints has become a prominen

safetyarxiv-cs-lg
4 May 2026
Safety

Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under Uncertainty

DGX agent

arXiv:2604.27193v1 Announce Type: new Abstract: Automatic Emergency Braking (AEB) systems represent a safety-critical national interest, with the National Highway Traffic Safety Administration (NHTSA)

safetyarxiv-cs-ro
1 May 2026
Safety

Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems

DGX agent

arXiv:2604.27358v1 Announce Type: new Abstract: As large language model (LLM) agents are deployed in high-stakes environments, the question of how safely to delegate subtasks to specialized sub-agents

safetyarxiv-cs-ai
1 May 2026
Model Releases

Intent Laundering: AI Safety Datasets Are Not What They Seem

DGX agent

arXiv:2602.16729v3 Announce Type: replace-cross Abstract: We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isol

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging

DGX agent

arXiv:2503.17239v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) is a common practice to adapt generalist models to specialized domains. However, recent studies show

safetyarxiv-cs-ai
24 Apr 2026
Model Releases

Secure LLM Fine-Tuning via Safety-Aware Probing

DGX agent

arXiv:2505.16737v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable success across many applications, but their ability to generate harmful content raises s

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

DGX agent

arXiv:2604.19559v1 Announce Type: new Abstract: Construction workers are highly vulnerable to heat stress, yet tools that translate real-time physiological data into actionable safety intelligence rem

safetyarxiv-cs-ai
22 Apr 2026
Safety

Safety-Critical Contextual Control via Online Riemannian Optimization with World Models

DGX agent

arXiv:2604.19639v1 Announce Type: cross Abstract: Modern world models are becoming too complex to admit explicit dynamical descriptions. We study safety-critical contextual control, where a Planner mu

safetyarxiv-cs-ai
22 Apr 2026
Model Releases

Guardrails in Logit Space: Safety Token Regularization for LLM Alignment

DGX agent

arXiv:2604.17210v1 Announce Type: new Abstract: Fine-tuning well-aligned large language models (LLMs) on new domains often degrades their safety alignment, even when using benign datasets. Existing sa

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification

DGX agent

arXiv:2604.17391v1 Announce Type: cross Abstract: RISC-V is emerging as a viable platform for automotive-grade embedded computing, with recent ISO 26262 ASIL-D certifications demonstrating readiness f

safetyarxiv-cs-lg
21 Apr 2026
Safety

Why Agents Compromise Safety Under Pressure

DGX agent

arXiv:2603.14975v2 Announce Type: replace-cross Abstract: Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering

safetyarxiv-cs-cl
21 Apr 2026
Safety

Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety

DGX agent

arXiv:2604.13101v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into aviation safety decision-making represents a significant technological advancement, yet their sta

safetyarxiv-cs-ai
17 Apr 2026
Safety

Injecting Hallucinations in Autonomous Vehicles: A Component-Agnostic Safety Evaluation Framework

DGX agent

arXiv:2510.07749v2 Announce Type: replace Abstract: Perception failures in autonomous vehicles (AV) remain a major safety concern because they are the basis for many accidents. To study how these fail

safetyarxiv-cs-ro
12 Aug 2026
Model Releases

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

DGX agent

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

DGX agent

arXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech

model-releasesarxiv-cs-ai
10 Aug 2026
Local Ai

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

DGX agent

arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward wh

local-aiarxiv-cs-ai
5 Aug 2026
Safety

Toward Certified Functional Safety for Industrial Humanoid Robots: The Fail-Passive Gap and a Feasibility Study

DGX agent

arXiv:2608.02809v1 Announce Type: new Abstract: Industrial humanoid robots are constrained less by locomotion or manipulation capability than by the immaturity of functional safety certification for l

safetyarxiv-cs-ro
5 Aug 2026
Safety

Towards General Language-Conditioned Latent Safety Filters

DGX agent

arXiv:2608.00315v1 Announce Type: cross Abstract: Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified

safetyarxiv-cs-lg
4 Aug 2026
Model Releases

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

DGX agent

arXiv:2607.22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation,

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

DGX agent

arXiv:2510.03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastru

safetyarxiv-cs-ai
28 Jul 2026
Safety

Nvidia forms the Open Secure AI Alliance, a coalition including CrowdStrike, Hugging Face, and Dell to develop and share tools for AI safety and cybersecurity (Jaspreet Singh/Reuters)

DGX agent

Jaspreet Singh / Reuters: Nvidia forms the Open Secure AI Alliance, a coalition including CrowdStrike, Hugging Face, and Dell to develop and share tools for AI safety and cybersecurity — Nvidia (NVDA.

safetytechmeme
27 Jul 2026
← Previous
123456…297
Next →