AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

DGX agent

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal

model-releasesarxiv-cs-cl
4 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Connected Dependability Cage: Run-Time Function and Anomaly Monitoring for the Development and Operation of Safe Automated Vehicles

DGX agent

arXiv:2604.27728v1 Announce Type: new Abstract: The advancement of automated vehicles introduces complex safety challenges, particularly in dynamic and unpredictable environments where AI-enabled perc

safetyarxiv-cs-ro
1 May 2026
Safety

Culturally Aware GenAI Risks for Youth: Perspectives from Youth, Parents, and Teachers in a Non-Western Context

DGX agent

arXiv:2604.26494v1 Announce Type: cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youth's safety i

safetyarxiv-cs-ai
30 Apr 2026
Safety

Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

DGX agent

arXiv:2604.22260v1 Announce Type: cross Abstract: Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While rece

safetyarxiv-cs-ai
27 Apr 2026
Safety

Atomic Decision Boundaries: A Structural Requirement for Guaranteeing Execution-Time Admissibility in Autonomous Systems

DGX agent

arXiv:2604.17511v2 Announce Type: replace-cross Abstract: Autonomous systems increasingly execute actions that directly modify shared state, creating an urgent need for precise control over which tran

safetyarxiv-cs-ai
23 Apr 2026
Safety

Stochastic Barrier Certificates in the Presence of Dynamic Obstacles

DGX agent

arXiv:2604.20208v1 Announce Type: new Abstract: Safety of stochastic dynamic systems in environments with dynamic obstacles is studied in this paper through the lens of stochastic barrier functions. W

safetyarxiv-cs-ro
23 Apr 2026
Safety

Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving

DGX agent

arXiv:2512.16055v2 Announce Type: replace Abstract: Safety-critical corner cases, difficult to collect in the real world, are crucial for evaluating end-to-end autonomous driving. Adversarial interact

safetyarxiv-cs-cv
21 Apr 2026
Safety

J-PARSE: Jacobian-based Projection Algorithm for Resolving Singularities Effectively in Inverse Kinematic Control of Serial Manipulators

DGX agent

arXiv:2505.00306v5 Announce Type: replace Abstract: J-PARSE is an algorithm for smooth first-order inverse kinematic control of a serial manipulator near kinematic singularities. The commanded end-eff

safetyarxiv-cs-ro
21 Apr 2026
Safety

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

DGX agent

arXiv:2604.17769v1 Announce Type: new Abstract: Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-e

safetyarxiv-cs-cl
21 Apr 2026
Safety

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

DGX agent

arXiv:2604.16606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains

safetyarxiv-cs-lg
21 Apr 2026
Safety

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

DGX agent

arXiv:2601.04740v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risk

safetyarxiv-cs-cl
21 Apr 2026
Safety

FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

DGX agent

arXiv:2604.15488v1 Announce Type: cross Abstract: Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offer

safetyarxiv-cs-ai
20 Apr 2026
Model Releases

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

DGX agent

arXiv:2505.24672v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploit

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Drowsiness-Aware Adaptive Autonomous Braking System based on Deep Reinforcement Learning for Enhanced Road Safety

DGX agent

arXiv:2604.13878v1 Announce Type: new Abstract: Driver drowsiness significantly impairs the ability to accurately judge safe braking distances and is estimated to contribute to 10%-20% of road acciden

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack

DGX agent

arXiv:2509.25843v2 Announce Type: replace Abstract: Large language models (LLMs), despite being safety-aligned, exhibit brittle refusal behaviors that can be circumvented by simple linguistic changes.

safetyarxiv-cs-ai
15 Apr 2026
Model Releases

EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems

DGX agent

arXiv:2604.11174v1 Announce Type: cross Abstract: Recent progress in embodied AI has produced a growing ecosystem of robot policies, foundation models, and modular runtimes. However, current evaluatio

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures

DGX agent

arXiv:2604.07709v1 Announce Type: cross Abstract: Ask a frontier model how to taper six milligrams of alprazolam (psychiatrist retired, ten days of pills left, abrupt cessation causes seizures) and it

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

KEO: Knowledge Extraction on OMIn via Knowledge Graphs and RAG for Safety-Critical Aviation Maintenance

DGX agent

arXiv:2510.05524v2 Announce Type: replace Abstract: We present Knowledge Extraction on OMIn (KEO), a domain-specific knowledge extraction and reasoning framework with large language models (LLMs) in s

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

DGX agent

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

DGX agent

arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evalu

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Do LLMs Know Their Vulnerable Scenarios?

DGX agent

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

DGX agent

arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Exist

safetyarxiv-cs-cl
27 Jul 2026
Safety

CommandLM: Data driven behavior level descriptor for ego vehicles

DGX agent

arXiv:2607.22078v1 Announce Type: new Abstract: As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decision-making is essential for safety, trust, and regul

safetyarxiv-cs-cv
27 Jul 2026
Safety

End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

DGX agent

arXiv:2607.20674v1 Announce Type: new Abstract: We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier function

safetyarxiv-cs-lg
24 Jul 2026
Safety

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

DGX agent

arXiv:2607.07601v1 Announce Type: cross Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize co

safetyarxiv-cs-ai
9 Jul 2026
Safety

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

DGX agent

arXiv:2607.02569v1 Announce Type: new Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using

safetyarxiv-cs-cv
7 Jul 2026
Safety

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

DGX agent

arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained fil

safetyarxiv-cs-cl
3 Jul 2026
Safety

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

DGX agent

arXiv:2606.23671v2 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and e

safetyarxiv-cs-cl
30 Jun 2026
Safety

In-Vehicle Digital Twin-Based Collision Warning Framework with Sybil Attack Detection

DGX agent

arXiv:2606.28625v1 Announce Type: cross Abstract: Connected Vehicles (CVs) rely extensively on communication technologies to enable data-driven predictive analyses for enhancing performance and safety

safetyarxiv-cs-lg
30 Jun 2026
Safety

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

DGX agent

arXiv:2606.12342v1 Announce Type: cross Abstract: Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language.

safetyarxiv-cs-ai
11 Jun 2026
Safety

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment

DGX agent

arXiv:2510.03520v2 Announce Type: replace-cross Abstract: Ensuring safety is a foundational requirement for large language models (LLMs). Achieving an appropriate balance between enhancing the utility

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

DGX agent

arXiv:2606.05660v1 Announce Type: new Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety t

model-releasesarxiv-cs-ro
5 Jun 2026
Model Releases

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format

DGX agent

arXiv:2509.02655v3 Announce Type: replace-cross Abstract: Many AI alignment discussions of 'runaway optimisation' focus on RL agents: unbounded utility maximisers that over-optimise a proxy objective

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Constitutional On-Policy Safe Distillation

DGX agent

arXiv:2606.03089v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to prov

safetyarxiv-cs-ai
3 Jun 2026
Safety

COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

DGX agent

arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf

safetyarxiv-cs-ai
1 Jun 2026
Safety

Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment

DGX agent

arXiv:2605.27659v1 Announce Type: cross Abstract: Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicle

safetyarxiv-cs-ai
28 May 2026
Safety

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty

DGX agent

arXiv:2605.26627v1 Announce Type: cross Abstract: Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising

safetyarxiv-cs-ro
27 May 2026
Safety

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

DGX agent

arXiv:2605.26158v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshol

safetyarxiv-cs-ai
27 May 2026
Safety

A Formal gatekeeper Framework for Safe Dual Control with Active Exploration

DGX agent

arXiv:2510.06351v2 Announce Type: replace Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations

safetyarxiv-cs-ro
26 May 2026
Safety

KG-ASG: Collision-Knowledge-Guided Closed-Loop Adversarial Scenario Generation With Primary-Support Attribution

DGX agent

arXiv:2605.18895v1 Announce Type: cross Abstract: Safety validation of autonomous driving systems requires high-risk scenario coverage, clear collision semantics, executable trajectories, and attribut

safetyarxiv-cs-ai
20 May 2026
Safety

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

DGX agent

arXiv:2605.19524v1 Announce Type: cross Abstract: End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models

safetyarxiv-cs-cv
20 May 2026
Model Releases

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion

DGX agent

arXiv:2605.11679v2 Announce Type: replace Abstract: In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. S

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

RISED: A Pre-Deployment Safety Evaluation Framework for Clinical AI Decision-Support Systems

DGX agent

arXiv:2605.12895v1 Announce Type: cross Abstract: Aggregate accuracy metrics dominate the evaluation of clinical AI decision-support systems but do not detect deployment-phase failures of input reliab

model-releasesarxiv-cs-ai
14 May 2026
Safety

Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware

DGX agent

arXiv:2506.00982v3 Announce Type: replace Abstract: Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, t

safetyarxiv-cs-ro
14 May 2026
Safety

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

DGX agent

arXiv:2605.08277v1 Announce Type: cross Abstract: Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demon

safetyarxiv-cs-ai
12 May 2026
Model Releases

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

DGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study

DGX agent

arXiv:2605.07422v1 Announce Type: cross Abstract: Qualitative analysis plays a pivotal role in understanding the human and social aspects of software engineering. However, it remains a demanding proce

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture

DGX agent

arXiv:2605.01567v1 Announce Type: cross Abstract: Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineeri

model-releasesarxiv-cs-cl
5 May 2026
← Previous
1…1415161718…255
Next →