AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Model Releases

Quality Is Not a Safety Proxy Under Quantization

DGX agent

arXiv:2606.10154v1 Announce Type: new Abstract: Quantized checkpoints are often screened first with quality metrics and only later, if at all, with direct safety tests. This paper audits that shortcut

model-releasesarxiv-cs-lg
10 Jun 2026
Safety

Impedance MPC for Physical Human-Robot Interaction: Predictive Disturbance Rejection with Joint-Limit Safety

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2606.08281v1 Announce Type: new Abstract: Physical human-robot interaction (pHRI) demands simultaneous trajectory accuracy and compliant safety under unplanned contact. Classical impedance contr

safetyarxiv-cs-ro
9 Jun 2026
Safety

Major improvement in safety statistics with FSD turned on in the Netherlands!

DGX agent

Elon Musk posted on X claiming that Tesla's Full Self-Driving (FSD) system resulted in significant improvements to vehicle safety statistics in the Netherlands. The post suggests FSD activation correl

safetyelon-musk--x
9 Jun 2026
Safety

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation

DGX agent

arXiv:2606.08414v1 Announce Type: cross Abstract: Diffusion policies have achieved remarkable success in robotic manipulation, yet they often fail to satisfy strict physical constraints required for s

safetyarxiv-cs-ai
9 Jun 2026
Safety

Building Pakistan Notice Helper: A Small AI Tool for a Very Local Safety Problem

DGX agent

Building Pakistan Notice Helper is a small AI tool designed to address a localized safety issue in Pakistan by helping users understand and process building-related notices. Developed as a hackathon p

safetyhugging-face
8 Jun 2026
Safety

Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety

DGX agent

arXiv:2606.05461v1 Announce Type: new Abstract: Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantifi

safetyarxiv-cs-ai
6 Jun 2026
Model Releases

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

DGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Meta expands Teen Accounts safety features to limit harmful content on Instagram, Facebook, and Messenger, including on nutrition, weight lifting, and anxiety (Eli Tan/New York Times)

DGX agent

Eli Tan / New York Times: Meta expands Teen Accounts safety features to limit harmful content on Instagram, Facebook, and Messenger, including on nutrition, weight lifting, and anxiety — The changes,

safetytechmeme
2 Jun 2026
Safety

Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

DGX agent

arXiv:2605.30368v1 Announce Type: cross Abstract: Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SS

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

DGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

model-releasesarxiv-cs-cl
29 May 2026
Safety

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

DGX agent

arXiv:2605.28664v1 Announce Type: cross Abstract: Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are sca

safetyarxiv-cs-cl
28 May 2026
Model Releases

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

DGX agent

arXiv:2601.01627v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety

model-releasesarxiv-cs-ai
28 May 2026
Safety

Safety-Critical Adaptive Impedance Control via Nonsmooth Control Barrier Functions under State and Input Constraints

DGX agent

arXiv:2605.28367v1 Announce Type: new Abstract: Safe physical interaction is critical for deploying robotic manipulators in human-robot interaction and contact-rich tasks, where uncertainty, external

safetyarxiv-cs-ro
28 May 2026
Safety

CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment

DGX agent

arXiv:2603.07211v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes tra

safetyarxiv-cs-lg
27 May 2026
Safety

Constrained Meta Reinforcement Learning with Provable Test-Time Safety

DGX agent

arXiv:2601.21845v2 Announce Type: replace Abstract: Meta reinforcement learning (RL) allows agents to leverage experience across a distribution of tasks on which the agent can train at will, enabling

safetyarxiv-cs-lg
27 May 2026
Safety

Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection

DGX agent

arXiv:2605.27155v1 Announce Type: cross Abstract: Testing object detectors in safety-critical domains requires semantically meaningful probes beyond pixel-level corruptions. We present SemProbe, a too

safetyarxiv-cs-ai
27 May 2026
Safety

CarlaNCAP: A Framework for Quantifying the Safety of Vulnerable Road Users in Infrastructure-Assisted Collective Perception Using EuroNCAP Scenarios

DGX agent

arXiv:2512.11551v2 Announce Type: replace Abstract: The growing number of road users has significantly increased the risk of accidents in recent years. Vulnerable Road Users (VRUs) are particularly at

safetyarxiv-cs-ro
25 May 2026
Local Ai

Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries

DGX agent

arXiv:2605.21712v1 Announce Type: new Abstract: Transportation safety analysis requires integrating crash records, roadway attributes, and geospatial data through GIS-based workflows, but access remai

local-aiarxiv-cs-cl
22 May 2026
Safety

Anomaly-Informed Confidence Calibration for Vision-Based Safety Prediction

DGX agent

arXiv:2605.21109v1 Announce Type: new Abstract: Reliable confidence estimates are important for safely deploying vision-based controllers in autonomous racing, where safety predictions must be derived

safetyarxiv-cs-ro
21 May 2026
Model Releases

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry

DGX agent

arXiv:2605.20241v1 Announce Type: cross Abstract: Prompt-level safety probes for large language models use hidden-state representations to separate safe from unsafe prompts, but strong average detecti

model-releasesarxiv-cs-cl
21 May 2026
Safety

Time-To-Reach Separation and Safety Filtering for Safe, Fair, and Efficient Multi-Agent Coordination

DGX agent

arXiv:2605.20625v1 Announce Type: cross Abstract: Advanced Air Mobility (AAM) operations are expected to significantly increase aerial traffic in urban airspace, requiring autonomous traffic managemen

safetyarxiv-cs-ro
21 May 2026
Safety

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

DGX agent

arXiv:2605.17413v1 Announce Type: cross Abstract: Safety-aligned language models often refuse cybersecurity requests whose wording resembles misuse, even when the task is authorized and defensive. Thi

safetyarxiv-cs-ai
19 May 2026
Safety

LLM-Safety Evaluations Lack Robustness

DGX agent

arXiv:2503.02574v2 Announce Type: replace-cross Abstract: In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of n

safetyarxiv-cs-ai
19 May 2026
Safety

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

DGX agent

arXiv:2603.18178v2 Announce Type: replace-cross Abstract: The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-co

safetyarxiv-cs-ai
19 May 2026
Model Releases

IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia

DGX agent

arXiv:2603.17915v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are deployed in multilingual settings, their safety behavior in culturally diverse, low-resource languages rem

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

DGX agent

arXiv:2509.26100v2 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, exis

model-releasesarxiv-cs-ai
15 May 2026
Safety

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

DGX agent

arXiv:2605.15164v1 Announce Type: cross Abstract: This position paper argues that behavioural assurance, even when carefully designed, is being asked to carry safety claims it cannot verify. AI govern

safetyarxiv-cs-ai
15 May 2026
Model Releases

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

DGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

model-releasesarxiv-cs-ai
15 May 2026
Safety

Before the Last Token: Diagnosing Final-Token Safety Probe Failures

DGX agent

arXiv:2605.12726v1 Announce Type: new Abstract: Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompts can contain probe-visible unsafe evidence distribute

safetyarxiv-cs-lg
14 May 2026
Safety

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

DGX agent

arXiv:2602.07668v2 Announce Type: replace Abstract: The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state

safetyarxiv-cs-cv
13 May 2026
Safety

Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

DGX agent

arXiv:2605.11769v1 Announce Type: new Abstract: Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. Whi

safetyarxiv-cs-cl
13 May 2026
Safety

Containment Verification: AI Safety Guarantees Independent of Alignment

DGX agent

arXiv:2605.09045v1 Announce Type: new Abstract: Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain

safetyarxiv-cs-ai
12 May 2026
Safety

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

DGX agent

arXiv:2605.02900v1 Announce Type: cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, saf

safetyarxiv-cs-cv
6 May 2026
Safety

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection

DGX agent

arXiv:2509.00673v2 Announce Type: replace Abstract: We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety

safetyarxiv-cs-cl
5 May 2026
Safety

TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks

DGX agent

arXiv:2605.01761v1 Announce Type: new Abstract: Text-to-Video (T2V) models have demonstrated remarkable capability in generating temporally coherent videos from natural language prompts, yet they also

safetyarxiv-cs-cv
5 May 2026
Safety

A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety

DGX agent

arXiv:2505.17342v2 Announce Type: replace Abstract: Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and

safetyarxiv-cs-lg
30 Apr 2026
Safety

From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Model

DGX agent

arXiv:2604.26052v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes such as attack success rate, refusal rate, or harmful/not-harmful re

safetyarxiv-cs-cl
30 Apr 2026
Model Releases

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

DGX agent

arXiv:2601.18491v2 Announce Type: replace Abstract: The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current gua

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

Reasoning Structure Matters for Safety Alignment of Reasoning Models

DGX agent

arXiv:2604.18946v1 Announce Type: new Abstract: Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This

safetyarxiv-cs-ai
22 Apr 2026
Safety

Vision-Based Human Awareness Estimation for Enhanced Safety and Efficiency of AMRs in Industrial Warehouses

DGX agent

arXiv:2604.18627v1 Announce Type: new Abstract: Ensuring human safety is of paramount importance in warehouse environments that feature mixed traffic of human workers and autonomous mobile robots (AMR

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains

DGX agent

arXiv:2604.19001v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm

model-releasesarxiv-cs-cl
22 Apr 2026
Safety

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

DGX agent

arXiv:2604.16424v1 Announce Type: cross Abstract: State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

DGX agent

arXiv:2601.04043v2 Announce Type: replace Abstract: As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

DGX agent

arXiv:2509.18127v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, un

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies

DGX agent

arXiv:2604.09189v1 Announce Type: cross Abstract: LLMs internalize safety policies through RLHF, yet these policies are never formally specified and remain difficult to inspect. Existing benchmarks ev

safetyarxiv-cs-ai
13 Apr 2026
Safety

OpenKedge: Governing Agentic Mutation with Execution-Bound Safety and Evidence Chains

DGX agent

arXiv:2604.08601v1 Announce Type: new Abstract: The rise of autonomous AI agents exposes a fundamental flaw in API-centric architectures: probabilistic systems directly execute state mutations without

safetyarxiv-cs-ai
13 Apr 2026
Safety

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

DGX agent

arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis

safetyarxiv-cs-cl
12 Aug 2026
← Previous
1…678910…297
Next →