AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
14 Apr 2026

EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems

Model ReleasesDGX agent

arXiv:2604.11174v1 Announce Type: cross Abstract: Recent progress in embodied AI has produced a growing ecosystem of robot policies, foundation models, and modular runtimes. However, current evaluatio

Closed-Form Concept Erasure via Double Projections

SafetyDGX agent

arXiv:2604.10032v1 Announce Type: cross Abstract: While modern generative models such as diffusion-based architectures have enabled impressive creative capabilities, they also raise important safety a

COSMIK-MPPI: Scaling Constrained Model Predictive Control to Collision Avoidance in Close-Proximity Dynamic Human Environments

SafetyDGX agent

arXiv:2604.10358v1 Announce Type: new Abstract: Ensuring safe physical interaction between torque-controlled manipulators and humans is essential for deploying robots in everyday environments. Model P

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

SafetyDGX agent

arXiv:2604.10326v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite ad

Learning to Test: Physics-Informed Representation for Dynamical Instability Detection

SafetyDGX agent

arXiv:2604.10967v1 Announce Type: new Abstract: Many safety-critical scientific and engineering systems evolve according to differential-algebraic equations (DAEs), where dynamical behavior is constra

Optimization-Guided Diffusion for Interactive Scene Generation

SafetyDGX agent

arXiv:2512.07661v3 Announce Type: replace Abstract: Realistic and diverse multi-agent driving scenes are crucial for evaluating autonomous vehicles, but safety-critical events which are essential for

PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk

SafetyDGX agent

arXiv:2604.11070v1 Announce Type: new Abstract: Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines

Safe Human-to-Humanoid Motion Imitation Using Control Barrier Functions

SafetyDGX agent

arXiv:2604.11447v1 Announce Type: new Abstract: Ensuring operational safety is critical for human-to-humanoid motion imitation. This paper presents a vision-based framework that enables a humanoid rob

10 Apr 2026

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures

Model ReleasesDGX agent

arXiv:2604.07709v1 Announce Type: cross Abstract: Ask a frontier model how to taper six milligrams of alprazolam (psychiatrist retired, ten days of pills left, abrupt cessation causes seizures) and it

KEO: Knowledge Extraction on OMIn via Knowledge Graphs and RAG for Safety-Critical Aviation Maintenance

Model ReleasesDGX agent

arXiv:2510.05524v2 Announce Type: replace Abstract: We present Knowledge Extraction on OMIn (KEO), a domain-specific knowledge extraction and reasoning framework with large language models (LLMs) in s

11 Aug 2026

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

Model ReleasesDGX agent

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-

5 Aug 2026

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

Model ReleasesDGX agent

arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evalu

4 Aug 2026

I added a verify-before-load safety check for Ollama models

Local AiDGX agent

I maintain llm-checker, and I’ve added structural model-file validation for Ollama. Ollama stores downloaded models as local blobs. If one is truncated, malformed, or has invalid internal offsets, you

28 Jul 2026

Do LLMs Know Their Vulnerable Scenarios?

Model ReleasesDGX agent

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa

27 Jul 2026

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

SafetyDGX agent

arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Exist

CommandLM: Data driven behavior level descriptor for ego vehicles

SafetyDGX agent

arXiv:2607.22078v1 Announce Type: new Abstract: As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decision-making is essential for safety, trust, and regul

25 Jul 2026

but they won’t.

SafetyDGX agent

but they won’t. AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better

24 Jul 2026

End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

SafetyDGX agent

arXiv:2607.20674v1 Announce Type: new Abstract: We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier function

9 Jul 2026

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

SafetyDGX agent

arXiv:2607.07601v1 Announce Type: cross Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize co

7 Jul 2026

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

SafetyDGX agent

arXiv:2607.02569v1 Announce Type: new Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using

3 Jul 2026

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

SafetyDGX agent

arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained fil

30 Jun 2026

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

SafetyDGX agent

arXiv:2606.23671v2 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and e

In-Vehicle Digital Twin-Based Collision Warning Framework with Sybil Attack Detection

SafetyDGX agent

arXiv:2606.28625v1 Announce Type: cross Abstract: Connected Vehicles (CVs) rely extensively on communication technologies to enable data-driven predictive analyses for enhancing performance and safety

11 Jun 2026

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

SafetyDGX agent

arXiv:2606.12342v1 Announce Type: cross Abstract: Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment

SafetyDGX agent

arXiv:2510.03520v2 Announce Type: replace-cross Abstract: Ensuring safety is a foundational requirement for large language models (LLMs). Achieving an appropriate balance between enhancing the utility

5 Jun 2026

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.05660v1 Announce Type: new Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety t

4 Jun 2026

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format

Model ReleasesDGX agent

arXiv:2509.02655v3 Announce Type: replace-cross Abstract: Many AI alignment discussions of 'runaway optimisation' focus on RL agents: unbounded utility maximisers that over-optimise a proxy objective

3 Jun 2026

Constitutional On-Policy Safe Distillation

SafetyDGX agent

arXiv:2606.03089v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to prov

1 Jun 2026

COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

SafetyDGX agent

arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf

28 May 2026

Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment

SafetyDGX agent

arXiv:2605.27659v1 Announce Type: cross Abstract: Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicle

27 May 2026

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty

SafetyDGX agent

arXiv:2605.26627v1 Announce Type: cross Abstract: Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

SafetyDGX agent

arXiv:2605.26158v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshol

26 May 2026

A Formal gatekeeper Framework for Safe Dual Control with Active Exploration

SafetyDGX agent

arXiv:2510.06351v2 Announce Type: replace Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations

20 May 2026

KG-ASG: Collision-Knowledge-Guided Closed-Loop Adversarial Scenario Generation With Primary-Support Attribution

SafetyDGX agent

arXiv:2605.18895v1 Announce Type: cross Abstract: Safety validation of autonomous driving systems requires high-risk scenario coverage, clear collision semantics, executable trajectories, and attribut

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

SafetyDGX agent

arXiv:2605.19524v1 Announce Type: cross Abstract: End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models

14 May 2026

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion

Model ReleasesDGX agent

arXiv:2605.11679v2 Announce Type: replace Abstract: In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. S

RISED: A Pre-Deployment Safety Evaluation Framework for Clinical AI Decision-Support Systems

Model ReleasesDGX agent

arXiv:2605.12895v1 Announce Type: cross Abstract: Aggregate accuracy metrics dominate the evaluation of clinical AI decision-support systems but do not detect deployment-phase failures of input reliab

Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware

SafetyDGX agent

arXiv:2506.00982v3 Announce Type: replace Abstract: Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, t

12 May 2026

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

SafetyDGX agent

arXiv:2605.08277v1 Announce Type: cross Abstract: Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demon

11 May 2026

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

Model ReleasesDGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study

Model ReleasesDGX agent

arXiv:2605.07422v1 Announce Type: cross Abstract: Qualitative analysis plays a pivotal role in understanding the human and social aspects of software engineering. However, it remains a demanding proce

5 May 2026

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture

Model ReleasesDGX agent

arXiv:2605.01567v1 Announce Type: cross Abstract: Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineeri

From Concept to Capability: Evaluating 3D Gaussian Splatting for Synthetic Scene Editing in Autonomous Driving

SafetyDGX agent

arXiv:2605.01995v1 Announce Type: new Abstract: The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while oper

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data

SafetyDGX agent

arXiv:2605.01356v1 Announce Type: new Abstract: Learning constraint-satisfying policies from offline data without risky online interaction is crucial for safety-critical decision making. Conventional

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

SafetyDGX agent

arXiv:2511.10586v2 Announce Type: replace-cross Abstract: Safe planning of an autonomous agent in interactive environments -- such as the control of a self-driving vehicle among pedestrians -- poses a

29 Apr 2026

ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution

SafetyDGX agent

arXiv:2604.25329v1 Announce Type: new Abstract: End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, a

28 Apr 2026

Computer Vision-Based Early Detection of Container Loss at Sea

SafetyDGX agent

arXiv:2604.24193v1 Announce Type: new Abstract: Containerised shipping underpins global trade, yet container loss at sea remains a persistent safety, environmental, and economic challenge. Despite com

INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety

SafetyDGX agent

arXiv:2604.23095v1 Announce Type: new Abstract: Indoor environments lack the spatial intelligence infrastructure that GPS provides outdoors; first responders arriving at unfamiliar buildings typically

SemML 2.0: Synthesizing Controllers for LTL

SafetyDGX agent

arXiv:2604.24102v1 Announce Type: new Abstract: Synthesizing a reactive system from specifications given in linear temporal logic (LTL) is a classical problem, finding its applications in safety-criti

Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture

SafetyDGX agent

arXiv:2604.23646v1 Announce Type: new Abstract: Recent evidence suggests that frontier AI systems can exhibit agentic misalignment, generating and executing harmful actions derived from internally con

27 Apr 2026

Estimating Tail Risks in Language Model Output Distributions

SafetyDGX agent

arXiv:2604.22167v1 Announce Type: cross Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increa

22 Apr 2026

HALO: Hybrid Auto-encoded Locomotion with Learned Latent Dynamics, Poincare Maps, and Regions of Attraction

SafetyDGX agent

arXiv:2604.18887v1 Announce Type: new Abstract: Reduced-order models are powerful for analyzing and controlling high-dimensional dynamical systems. Yet constructing these models for complex hybrid sys

21 Apr 2026

Heterogeneous Self-Play for Realistic Highway Traffic Simulation

SafetyDGX agent

arXiv:2604.16406v1 Announce Type: cross Abstract: Realistic highway simulation is critical for scalable safety evaluation of autonomous vehicles, particularly for interactions that are too rare to stu

17 Apr 2026

Towards Verified and Targeted Explanations through Formal Methods

SafetyDGX agent

arXiv:2604.14209v1 Announce Type: new Abstract: As deep neural networks are deployed in safety-critical domains such as autonomous driving and medical diagnosis, stakeholders need explanations that ar

16 Apr 2026

Empirical Prediction of Pedestrian Comfort in Mobile Robot Pedestrian Encounters

SafetyDGX agent

arXiv:2604.13677v1 Announce Type: new Abstract: Mobile robots joining public spaces like sidewalks must care for pedestrian comfort. Many studies consider pedestrians' objective safety, for example, b

Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions

Model ReleasesDGX agent

arXiv:2601.03173v2 Announce Type: replace Abstract: Time pressure critically influences risky maneuvers and crash proneness among powered two-wheeler riders, yet its prediction remains underexplored i

15 Apr 2026

Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors

SafetyDGX agent

arXiv:2604.12359v1 Announce Type: cross Abstract: Safety-aligned large language models (LLMs) are increasingly deployed in real-world pipelines, yet this deployment also enlarges the supply-chain atta

Safe-FedLLM: Delving into the Safety of Federated Large Language Models

Local AiDGX agent

arXiv:2601.07177v2 Announce Type: replace-cross Abstract: Federated learning (FL) addresses privacy and data-silo issues in the training of large language models (LLMs). Most prior work focuses on imp

Scalable Verification of Neural Control Barrier Functions Using Linear Bound Propagation

SafetyDGX agent

arXiv:2511.06341v2 Announce Type: replace Abstract: Control barrier functions (CBFs) are a popular tool for safety certification of nonlinear dynamical control systems. Recently, CBFs represented as n

13 Apr 2026

Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models

SafetyDGX agent

arXiv:2604.08557v1 Announce Type: cross Abstract: Diffusion-based language models (dLLMs) generate text by iteratively denoising masked token sequences. We show that their safety alignment rests on a

← Previous
1…1314151617…238
Next →