AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

DGX agent

arXiv:2605.26158v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshol

safetyarxiv-cs-ai
27 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

A Formal gatekeeper Framework for Safe Dual Control with Active Exploration

DGX agent

arXiv:2510.06351v2 Announce Type: replace Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations

safetyarxiv-cs-ro
26 May 2026
Safety

KG-ASG: Collision-Knowledge-Guided Closed-Loop Adversarial Scenario Generation With Primary-Support Attribution

DGX agent

arXiv:2605.18895v1 Announce Type: cross Abstract: Safety validation of autonomous driving systems requires high-risk scenario coverage, clear collision semantics, executable trajectories, and attribut

safetyarxiv-cs-ai
20 May 2026
Safety

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

DGX agent

arXiv:2605.19524v1 Announce Type: cross Abstract: End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models

safetyarxiv-cs-cv
20 May 2026
Model Releases

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion

DGX agent

arXiv:2605.11679v2 Announce Type: replace Abstract: In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. S

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

RISED: A Pre-Deployment Safety Evaluation Framework for Clinical AI Decision-Support Systems

DGX agent

arXiv:2605.12895v1 Announce Type: cross Abstract: Aggregate accuracy metrics dominate the evaluation of clinical AI decision-support systems but do not detect deployment-phase failures of input reliab

model-releasesarxiv-cs-ai
14 May 2026
Safety

Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware

DGX agent

arXiv:2506.00982v3 Announce Type: replace Abstract: Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, t

safetyarxiv-cs-ro
14 May 2026
Safety

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

DGX agent

arXiv:2605.08277v1 Announce Type: cross Abstract: Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demon

safetyarxiv-cs-ai
12 May 2026
Model Releases

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

DGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study

DGX agent

arXiv:2605.07422v1 Announce Type: cross Abstract: Qualitative analysis plays a pivotal role in understanding the human and social aspects of software engineering. However, it remains a demanding proce

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture

DGX agent

arXiv:2605.01567v1 Announce Type: cross Abstract: Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineeri

model-releasesarxiv-cs-cl
5 May 2026
Safety

From Concept to Capability: Evaluating 3D Gaussian Splatting for Synthetic Scene Editing in Autonomous Driving

DGX agent

arXiv:2605.01995v1 Announce Type: new Abstract: The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while oper

safetyarxiv-cs-cv
5 May 2026
Safety

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data

DGX agent

arXiv:2605.01356v1 Announce Type: new Abstract: Learning constraint-satisfying policies from offline data without risky online interaction is crucial for safety-critical decision making. Conventional

safetyarxiv-cs-lg
5 May 2026
Safety

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

DGX agent

arXiv:2511.10586v2 Announce Type: replace-cross Abstract: Safe planning of an autonomous agent in interactive environments -- such as the control of a self-driving vehicle among pedestrians -- poses a

safetyarxiv-cs-ro
5 May 2026
Safety

ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution

DGX agent

arXiv:2604.25329v1 Announce Type: new Abstract: End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, a

safetyarxiv-cs-ro
29 Apr 2026
Safety

Computer Vision-Based Early Detection of Container Loss at Sea

DGX agent

arXiv:2604.24193v1 Announce Type: new Abstract: Containerised shipping underpins global trade, yet container loss at sea remains a persistent safety, environmental, and economic challenge. Despite com

safetyarxiv-cs-cv
28 Apr 2026
Safety

INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety

DGX agent

arXiv:2604.23095v1 Announce Type: new Abstract: Indoor environments lack the spatial intelligence infrastructure that GPS provides outdoors; first responders arriving at unfamiliar buildings typically

safetyarxiv-cs-cv
28 Apr 2026
Safety

SemML 2.0: Synthesizing Controllers for LTL

DGX agent

arXiv:2604.24102v1 Announce Type: new Abstract: Synthesizing a reactive system from specifications given in linear temporal logic (LTL) is a classical problem, finding its applications in safety-criti

safetyarxiv-cs-ai
28 Apr 2026
Safety

Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture

DGX agent

arXiv:2604.23646v1 Announce Type: new Abstract: Recent evidence suggests that frontier AI systems can exhibit agentic misalignment, generating and executing harmful actions derived from internally con

safetyarxiv-cs-ai
28 Apr 2026
Safety

Estimating Tail Risks in Language Model Output Distributions

DGX agent

arXiv:2604.22167v1 Announce Type: cross Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increa

safetyarxiv-cs-ai
27 Apr 2026
Safety

HALO: Hybrid Auto-encoded Locomotion with Learned Latent Dynamics, Poincare Maps, and Regions of Attraction

DGX agent

arXiv:2604.18887v1 Announce Type: new Abstract: Reduced-order models are powerful for analyzing and controlling high-dimensional dynamical systems. Yet constructing these models for complex hybrid sys

safetyarxiv-cs-ro
22 Apr 2026
Safety

Heterogeneous Self-Play for Realistic Highway Traffic Simulation

DGX agent

arXiv:2604.16406v1 Announce Type: cross Abstract: Realistic highway simulation is critical for scalable safety evaluation of autonomous vehicles, particularly for interactions that are too rare to stu

safetyarxiv-cs-lg
21 Apr 2026
Safety

Towards Verified and Targeted Explanations through Formal Methods

DGX agent

arXiv:2604.14209v1 Announce Type: new Abstract: As deep neural networks are deployed in safety-critical domains such as autonomous driving and medical diagnosis, stakeholders need explanations that ar

safetyarxiv-cs-lg
17 Apr 2026
Safety

Empirical Prediction of Pedestrian Comfort in Mobile Robot Pedestrian Encounters

DGX agent

arXiv:2604.13677v1 Announce Type: new Abstract: Mobile robots joining public spaces like sidewalks must care for pedestrian comfort. Many studies consider pedestrians' objective safety, for example, b

safetyarxiv-cs-ro
16 Apr 2026
Model Releases

Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions

DGX agent

arXiv:2601.03173v2 Announce Type: replace Abstract: Time pressure critically influences risky maneuvers and crash proneness among powered two-wheeler riders, yet its prediction remains underexplored i

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors

DGX agent

arXiv:2604.12359v1 Announce Type: cross Abstract: Safety-aligned large language models (LLMs) are increasingly deployed in real-world pipelines, yet this deployment also enlarges the supply-chain atta

safetyarxiv-cs-cl
15 Apr 2026
Local Ai

Safe-FedLLM: Delving into the Safety of Federated Large Language Models

DGX agent

arXiv:2601.07177v2 Announce Type: replace-cross Abstract: Federated learning (FL) addresses privacy and data-silo issues in the training of large language models (LLMs). Most prior work focuses on imp

local-aiarxiv-cs-ai
15 Apr 2026
Safety

Scalable Verification of Neural Control Barrier Functions Using Linear Bound Propagation

DGX agent

arXiv:2511.06341v2 Announce Type: replace Abstract: Control barrier functions (CBFs) are a popular tool for safety certification of nonlinear dynamical control systems. Recently, CBFs represented as n

safetyarxiv-cs-lg
15 Apr 2026
Safety

Closed-Form Concept Erasure via Double Projections

DGX agent

arXiv:2604.10032v1 Announce Type: cross Abstract: While modern generative models such as diffusion-based architectures have enabled impressive creative capabilities, they also raise important safety a

safetyarxiv-cs-ai
14 Apr 2026
Safety

COSMIK-MPPI: Scaling Constrained Model Predictive Control to Collision Avoidance in Close-Proximity Dynamic Human Environments

DGX agent

arXiv:2604.10358v1 Announce Type: new Abstract: Ensuring safe physical interaction between torque-controlled manipulators and humans is essential for deploying robots in everyday environments. Model P

safetyarxiv-cs-ro
14 Apr 2026
Safety

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

DGX agent

arXiv:2604.10326v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite ad

safetyarxiv-cs-ai
14 Apr 2026
Safety

Learning to Test: Physics-Informed Representation for Dynamical Instability Detection

DGX agent

arXiv:2604.10967v1 Announce Type: new Abstract: Many safety-critical scientific and engineering systems evolve according to differential-algebraic equations (DAEs), where dynamical behavior is constra

safetyarxiv-cs-lg
14 Apr 2026
Safety

Optimization-Guided Diffusion for Interactive Scene Generation

DGX agent

arXiv:2512.07661v3 Announce Type: replace Abstract: Realistic and diverse multi-agent driving scenes are crucial for evaluating autonomous vehicles, but safety-critical events which are essential for

safetyarxiv-cs-cv
14 Apr 2026
Safety

PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk

DGX agent

arXiv:2604.11070v1 Announce Type: new Abstract: Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines

safetyarxiv-cs-ai
14 Apr 2026
Safety

Safe Human-to-Humanoid Motion Imitation Using Control Barrier Functions

DGX agent

arXiv:2604.11447v1 Announce Type: new Abstract: Ensuring operational safety is critical for human-to-humanoid motion imitation. This paper presents a vision-based framework that enables a humanoid rob

safetyarxiv-cs-ro
14 Apr 2026
Safety

Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models

DGX agent

arXiv:2604.08557v1 Announce Type: cross Abstract: Diffusion-based language models (dLLMs) generate text by iteratively denoising masked token sequences. We show that their safety alignment rests on a

safetyarxiv-cs-ai
13 Apr 2026
Safety

SafeMind: A Risk-Aware Differentiable Control Framework for Adaptive and Safe Quadruped Locomotion

DGX agent

arXiv:2604.09474v1 Announce Type: cross Abstract: Learning-based quadruped controllers achieve impressive agility but typically lack formal safety guarantees under model uncertainty, perception noise,

safetyarxiv-cs-ai
13 Apr 2026
Safety

A Giant-Step Baby-Step Classifier For Scalable and Real-Time Anomaly Detection In Industrial Control Systems and Water Treatment Systems

DGX agent

arXiv:2504.20906v4 Announce Type: replace-cross Abstract: The continuous monitoring of the interactions between cyber-physical components of any industrial control system (ICS) is required to secure a

safetyarxiv-cs-lg
10 Apr 2026
Safety

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

DGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

safetyarxiv-cs-ai
12 Aug 2026
Safety

Graph-Guided Safe Diffuser: Topological Graph Guidance for Safe Diffusion Planning

DGX agent

arXiv:2608.09484v1 Announce Type: new Abstract: Many diffusion-based planners enforce safety through inference-time guidance, but such interleaved trajectory deformations often degrade kinematic feasi

safetyarxiv-cs-ro
11 Aug 2026
Safety

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

DGX agent

arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Howev

safetyarxiv-cs-ai
11 Aug 2026
Safety

Forbidden Region Dynamic Active Constraints in Robot-Assisted Minimally Invasive Surgery

DGX agent

arXiv:2608.03010v1 Announce Type: new Abstract: In robot-assisted surgery, Forbidden Region Active Constraints (FRAC) represent a control strategy that helps maintain task safety by generating anisotr

safetyarxiv-cs-ro
5 Aug 2026
Safety

Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling

DGX agent

arXiv:2607.27511v1 Announce Type: new Abstract: Imitation learning has shown increasing promise for autonomous robotic surgery, yet safe deployment remains challenging due to the safety-critical natur

safetyarxiv-cs-ro
31 Jul 2026
Safety

Uncertainty quantification for trustworthy deep learning: Methods and measures

DGX agent

arXiv:2607.28248v1 Announce Type: cross Abstract: The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures

safetyarxiv-cs-lg
31 Jul 2026
Safety

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

DGX agent

arXiv:2607.26574v1 Announce Type: cross Abstract: Safety classifiers ('guards') are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning:

safetyarxiv-cs-lg
30 Jul 2026
Safety

Steering Instruction Hierarchies at Inference Time

DGX agent

arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override confl

safetyarxiv-cs-cl
30 Jul 2026
Safety

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, man…

DGX agent

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, many forefront members of the AI-Safety community, in their fer

safetyyann-lecun--x
30 Jul 2026
Safety

Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail

DGX agent

arXiv:2607.25664v1 Announce Type: new Abstract: Machine learning demand forecasts optimize statistical accuracy yet leave excess operational volatility that inflates safety stock and amplifies the Bul

safetyarxiv-cs-lg
29 Jul 2026
← Previous
1…1718192021…297
Next →