AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,230 results
Safety

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning

DGX agent

arXiv:2606.00160v1 Announce Type: cross Abstract: Large language models (LLMs) suffer from degraded safety capabilities even when fine-tuned with benign datasets. However, existing methods for identif

safetyarxiv-cs-ai
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

DGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

safetyarxiv-cs-ai
26 May 2026
Safety

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction

DGX agent

arXiv:2605.18104v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often fail to transfer safety capabilities learned in the text modality to semantically equivalent non-text inp

safetyarxiv-cs-ai
19 May 2026
Safety

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions

DGX agent

arXiv:2408.12935v4 Announce Type: replace Abstract: AI Safety is an emerging area of critical importance to the safe adoption and deployment of AI systems. With the rapid proliferation of AI and espec

safetyarxiv-cs-ai
14 May 2026
Safety

BSO: Safety Alignment Is Density Ratio Matching

DGX agent

arXiv:2605.12339v1 Announce Type: new Abstract: Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement lear

safetyarxiv-cs-lg
13 May 2026
Safety

Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning

DGX agent

arXiv:2602.17546v2 Announce Type: replace Abstract: Instruction-following language models are trained to be helpful and safe, yet their safety behavior can deteriorate under benign fine-tuning and wor

safetyarxiv-cs-cl
12 May 2026
Safety

Mental Health AI Safety Claims Must Preserve Temporal Evidence

DGX agent

arXiv:2605.08827v1 Announce Type: new Abstract: The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, o

safetyarxiv-cs-ai
12 May 2026
Safety

Multilingual Safety Alignment via Self-Distillation

DGX agent

arXiv:2605.02971v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain hig

safetyarxiv-cs-cl
6 May 2026
Safety

Online Safety Filter for Deformable Object Manipulation with Horizon Agnostic Neural Operators

DGX agent

arXiv:2605.01069v1 Announce Type: new Abstract: Safety critical control of robotic manipulation tasks involving deformable media such as fluids, cloth, and soft objects remains challenging because exi

safetyarxiv-cs-ro
5 May 2026
Safety

Grammar-Constrained Refinement of Safety Operational Rules Using Language in the Loop: What Could Go Wrong

DGX agent

arXiv:2604.23523v1 Announce Type: cross Abstract: Safety specifications in cyber-physical systems (CPS) capture the operational conditions the system must satisfy to operate safely within its intended

safetyarxiv-cs-ai
28 Apr 2026
Safety

SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs

DGX agent

arXiv:2604.20930v1 Announce Type: cross Abstract: Internal Safety Collapse (ISC) is a failure mode in which frontier LLMs, when executing legitimate professional tasks whose correct completion structu

safetyarxiv-cs-ai
24 Apr 2026
Safety

Continual Safety Alignment via Gradient-Based Sample Selection

DGX agent

arXiv:2604.17215v1 Announce Type: new Abstract: Large language models require continuous adaptation to new tasks while preserving safety alignment. However, fine-tuning on even benign data often compr

safetyarxiv-cs-lg
21 Apr 2026
Safety

On Safety Risks in Experience-Driven Self-Evolving Agents

DGX agent

arXiv:2604.16968v1 Announce Type: new Abstract: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self

safetyarxiv-cs-cl
21 Apr 2026
Safety

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints

DGX agent

arXiv:2604.12384v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) remains highly fragile during fine-tuning, where even benign adaptation can degrade pre-trained refusal

safetyarxiv-cs-ai
15 Apr 2026
Safety

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score…

DGX agent

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score of 100. This allows you to maintain a higher average safety

safetyelon-musk--x
14 Apr 2026
Local Ai

The Illusion of Cross-Lingual Safety in Low-Resource Languages

DGX agent

arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. How

local-aiarxiv-cs-cl
12 Aug 2026
Safety

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

DGX agent

arXiv:2608.07535v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding an

safetyarxiv-cs-ai
11 Aug 2026
Safety

Safety Cost of Steering Vectors Is Separable and Reducible

DGX agent

arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally comprom

safetyarxiv-cs-cl
11 Aug 2026
Safety

White House, AI firms keep safety framework talks private

DGX agent

The White House met with representatives from leading artificial intelligence companies today to discuss a safety framework for the government to review frontier models prior to launch, although there

safetysiliconangle
5 Aug 2026
Safety

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

DGX agent

arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, a

safetyarxiv-cs-lg
31 Jul 2026
Model Releases

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

DGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

model-releasesarxiv-cs-cl
28 Jul 2026
Safety

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

DGX agent

arXiv:2607.19913v1 Announce Type: new Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-orient

safetyarxiv-cs-ai
23 Jul 2026
Safety

Harnessing Textual Refusal Directions for Multimodal Safety

DGX agent

arXiv:2606.31876v1 Announce Type: new Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. B

safetyarxiv-cs-ai
1 Jul 2026
Safety

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion

DGX agent

arXiv:2606.25097v1 Announce Type: new Abstract: Speculative decoding accelerates inference by letting a draft model propose tokens for a target model to verify, raising a concrete safety question: at

safetyarxiv-cs-lg
25 Jun 2026
Safety

Towards a Bathroom-Centered Human-Building Digital Twin Framework for Indoor Safety Analysis

DGX agent

arXiv:2606.23292v2 Announce Type: replace-cross Abstract: Bathroom use is a critical safety challenge for older adults because wet surfaces, constrained layouts, limited support, and frequent posture

safetyarxiv-cs-ai
25 Jun 2026
Safety

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

DGX agent

arXiv:2606.25034v1 Announce Type: new Abstract: General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversaria

safetyarxiv-cs-cv
25 Jun 2026
Safety

Safe-SAGE: Social-Semantic Adaptive Guidance for Safe Engagement through Laplace-Modulated Poisson Safety Functions

DGX agent

arXiv:2603.05497v3 Announce Type: replace Abstract: Traditional safety-critical control methods, such as control barrier functions, suffer from semantic blindness, exhibiting the same behavior around

safetyarxiv-cs-ro
23 Jun 2026
Safety

Listening to the Workforce: Measuring Construction Worker Safety Attitudes from Social Media Discourse Using LLMs

DGX agent

arXiv:2606.04450v1 Announce Type: new Abstract: Worker safety attitudes are key determinants of whether protective practices are applied or bypassed on construction sites. Yet measuring them at scale

safetyarxiv-cs-cl
4 Jun 2026
Safety

Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections

DGX agent

arXiv:2605.27418v1 Announce Type: cross Abstract: The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transp

safetyarxiv-cs-ro
28 May 2026
Model Releases

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

DGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

model-releasesarxiv-cs-cl
28 May 2026
Safety

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

DGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

safetyarxiv-cs-ai
26 May 2026
Safety

Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays

DGX agent

arXiv:2605.23351v1 Announce Type: new Abstract: We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while

safetyarxiv-cs-lg
25 May 2026
Safety

Generating Realistic Safety-Critical Scenarios for Vehicle-Pedestrian Interactions

DGX agent

arXiv:2605.17229v1 Announce Type: new Abstract: Automated driving system deployment requires rigorous validation across safety-critical vehicle-pedestrian interactions, yet real-world datasets rarely

safetyarxiv-cs-ro
19 May 2026
Safety

Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents

DGX agent

arXiv:2605.17830v1 Announce Type: new Abstract: Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under

safetyarxiv-cs-ai
19 May 2026
Safety

SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation

DGX agent

arXiv:2601.18442v3 Announce Type: replace Abstract: Autonomous Vehicle (AV) requires rigorous testing in safety-critical scenarios for safety validation, yet its validation is hindered by the high cos

safetyarxiv-cs-ro
19 May 2026
Safety

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability

DGX agent

arXiv:2605.14246v1 Announce Type: cross Abstract: Many safety-critical control problems are modeled as risk-sensitive partially observable Markov decision processes, where the controller must make dec

safetyarxiv-cs-ai
15 May 2026
Safety

Selective Safety Steering via Value-Filtered Decoding

DGX agent

arXiv:2605.14746v1 Announce Type: new Abstract: While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of wor

safetyarxiv-cs-lg
15 May 2026
Safety

Internalizing Safety Understanding in Large Reasoning Models via Verification

DGX agent

arXiv:2605.08930v1 Announce Type: new Abstract: While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment para

safetyarxiv-cs-ai
12 May 2026
Safety

Shields to Guarantee Probabilistic Safety in MDPs

DGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

safetyarxiv-cs-ai
12 May 2026
Safety

A Closed-Form Dual-Barrier CBF Safety Filter for Holonomic Robots on Incrementally Built Occupancy Grid Maps

DGX agent

arXiv:2605.05182v1 Announce Type: new Abstract: We present a dual-barrier control barrier function (CBF) safety filter for real-time, safety-critical velocity control of holonomic robots operating in

safetyarxiv-cs-ro
7 May 2026
Safety

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

DGX agent

arXiv:2605.01147v1 Announce Type: new Abstract: As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety propertie

safetyarxiv-cs-ai
6 May 2026
Safety

Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics

DGX agent

arXiv:2604.26836v1 Announce Type: new Abstract: Predictive safety filters (PSFs) leverage model predictive control to enforce constraint satisfaction during deep reinforcement learning (RL) exploratio

safetyarxiv-cs-lg
30 Apr 2026
Safety

Unifying Runtime Monitoring Approaches for Safety-Critical Machine Learning: Application to Vision-Based Landing

DGX agent

arXiv:2604.26411v1 Announce Type: new Abstract: Runtime monitoring is essential to ensure the safety of ML applications in safety-critical domains. However, current research is fragmented, with indepe

safetyarxiv-cs-lg
30 Apr 2026
Safety

LLM-Augmented Traffic Signal Control with LSTM-Based Traffic State Prediction and Safety-Constrained Decision Support

DGX agent

arXiv:2604.23902v1 Announce Type: new Abstract: Traffic signal control is a critical task in intelligent transportation systems, yet conventional fixed-time and rule-based methods often struggle to ad

safetyarxiv-cs-ai
28 Apr 2026
Safety

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

DGX agent

arXiv:2604.24348v1 Announce Type: new Abstract: The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS

safetyarxiv-cs-cl
28 Apr 2026
Safety

TSAssistant: A Human-in-the-Loop Agentic Framework for Automated Target Safety Assessment

DGX agent

arXiv:2604.23938v1 Announce Type: new Abstract: Target Safety Assessment (TSA) requires systematic integration of heterogeneous evidence, including genetic, transcriptomic, target homology, pharmacolo

safetyarxiv-cs-cl
28 Apr 2026
Safety

Cat-DPO: Category-Adaptive Safety Alignment

DGX agent

arXiv:2604.17299v1 Announce Type: new Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusin

safetyarxiv-cs-cl
21 Apr 2026
Safety

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

DGX agent

arXiv:2604.17730v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging d

safetyarxiv-cs-cl
21 Apr 2026
← Previous
12345…297
Next →