AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Guiding Neuro-Symbolic Scenario Generation with Spatio-Temporal Logic

DGX agent

arXiv:2605.19038v1 Announce Type: cross Abstract: The rapid advancement of autonomous driving (AD) technologies has outpaced the development of robust safety evaluation methods. Conventional testing r

safetyarxiv-cs-lg
20 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails

DGX agent

arXiv:2605.17329v1 Announce Type: cross Abstract: Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, s

safetyarxiv-cs-ai
19 May 2026
Local Ai

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

DGX agent

arXiv:2605.14454v1 Announce Type: cross Abstract: As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line

local-aiarxiv-cs-cl
15 May 2026
Model Releases

Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks

DGX agent

arXiv:2605.14604v1 Announce Type: new Abstract: This position paper argues that effective tutoring requires corrective friction: surfacing misconceptions and challenging them supportively to drive con

model-releasesarxiv-cs-ai
15 May 2026
Safety

Metaphor Is Not All Attention Needs

DGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

safetyarxiv-cs-cl
13 May 2026
Safety

Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization

DGX agent

arXiv:2409.10310v3 Announce Type: replace Abstract: Ensuring safety and driving consistency is a significant challenge for autonomous vehicles operating in partially observed environments. This work i

safetyarxiv-cs-ro
12 May 2026
Local Ai

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

DGX agent

arXiv:2605.08878v1 Announce Type: cross Abstract: Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representa

local-aiarxiv-cs-ai
12 May 2026
Model Releases

GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access

DGX agent

Executive Summary Since our February 2026 report on AI-related threat activity, Google Threat Intelligence Group (GTIG) has continued to track a maturing transition from nascent AI-enabled operations

model-releasesgoogle-cloud-ai
11 May 2026
Safety

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

DGX agent

arXiv:2601.19487v2 Announce Type: replace Abstract: Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector

safetyarxiv-cs-lg
5 May 2026
Model Releases

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

DGX agent

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal

model-releasesarxiv-cs-cl
4 May 2026
Safety

Connected Dependability Cage: Run-Time Function and Anomaly Monitoring for the Development and Operation of Safe Automated Vehicles

DGX agent

arXiv:2604.27728v1 Announce Type: new Abstract: The advancement of automated vehicles introduces complex safety challenges, particularly in dynamic and unpredictable environments where AI-enabled perc

safetyarxiv-cs-ro
1 May 2026
Safety

Culturally Aware GenAI Risks for Youth: Perspectives from Youth, Parents, and Teachers in a Non-Western Context

DGX agent

arXiv:2604.26494v1 Announce Type: cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youth's safety i

safetyarxiv-cs-ai
30 Apr 2026
Safety

Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

DGX agent

arXiv:2604.22260v1 Announce Type: cross Abstract: Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While rece

safetyarxiv-cs-ai
27 Apr 2026
Safety

Atomic Decision Boundaries: A Structural Requirement for Guaranteeing Execution-Time Admissibility in Autonomous Systems

DGX agent

arXiv:2604.17511v2 Announce Type: replace-cross Abstract: Autonomous systems increasingly execute actions that directly modify shared state, creating an urgent need for precise control over which tran

safetyarxiv-cs-ai
23 Apr 2026
Safety

Stochastic Barrier Certificates in the Presence of Dynamic Obstacles

DGX agent

arXiv:2604.20208v1 Announce Type: new Abstract: Safety of stochastic dynamic systems in environments with dynamic obstacles is studied in this paper through the lens of stochastic barrier functions. W

safetyarxiv-cs-ro
23 Apr 2026
Safety

Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving

DGX agent

arXiv:2512.16055v2 Announce Type: replace Abstract: Safety-critical corner cases, difficult to collect in the real world, are crucial for evaluating end-to-end autonomous driving. Adversarial interact

safetyarxiv-cs-cv
21 Apr 2026
Safety

J-PARSE: Jacobian-based Projection Algorithm for Resolving Singularities Effectively in Inverse Kinematic Control of Serial Manipulators

DGX agent

arXiv:2505.00306v5 Announce Type: replace Abstract: J-PARSE is an algorithm for smooth first-order inverse kinematic control of a serial manipulator near kinematic singularities. The commanded end-eff

safetyarxiv-cs-ro
21 Apr 2026
Safety

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

DGX agent

arXiv:2604.17769v1 Announce Type: new Abstract: Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-e

safetyarxiv-cs-cl
21 Apr 2026
Safety

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

DGX agent

arXiv:2604.16606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains

safetyarxiv-cs-lg
21 Apr 2026
Safety

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

DGX agent

arXiv:2601.04740v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risk

safetyarxiv-cs-cl
21 Apr 2026
Safety

FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

DGX agent

arXiv:2604.15488v1 Announce Type: cross Abstract: Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offer

safetyarxiv-cs-ai
20 Apr 2026
Model Releases

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

DGX agent

arXiv:2505.24672v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploit

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Drowsiness-Aware Adaptive Autonomous Braking System based on Deep Reinforcement Learning for Enhanced Road Safety

DGX agent

arXiv:2604.13878v1 Announce Type: new Abstract: Driver drowsiness significantly impairs the ability to accurately judge safe braking distances and is estimated to contribute to 10%-20% of road acciden

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack

DGX agent

arXiv:2509.25843v2 Announce Type: replace Abstract: Large language models (LLMs), despite being safety-aligned, exhibit brittle refusal behaviors that can be circumvented by simple linguistic changes.

safetyarxiv-cs-ai
15 Apr 2026
Model Releases

EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems

DGX agent

arXiv:2604.11174v1 Announce Type: cross Abstract: Recent progress in embodied AI has produced a growing ecosystem of robot policies, foundation models, and modular runtimes. However, current evaluatio

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures

DGX agent

arXiv:2604.07709v1 Announce Type: cross Abstract: Ask a frontier model how to taper six milligrams of alprazolam (psychiatrist retired, ten days of pills left, abrupt cessation causes seizures) and it

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

KEO: Knowledge Extraction on OMIn via Knowledge Graphs and RAG for Safety-Critical Aviation Maintenance

DGX agent

arXiv:2510.05524v2 Announce Type: replace Abstract: We present Knowledge Extraction on OMIn (KEO), a domain-specific knowledge extraction and reasoning framework with large language models (LLMs) in s

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

DGX agent

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

DGX agent

arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evalu

model-releasesarxiv-cs-cv
5 Aug 2026
Local Ai

I added a verify-before-load safety check for Ollama models

DGX agent

I maintain llm-checker, and I’ve added structural model-file validation for Ollama. Ollama stores downloaded models as local blobs. If one is truncated, malformed, or has invalid internal offsets, you

local-air-ollama
4 Aug 2026
Model Releases

Do LLMs Know Their Vulnerable Scenarios?

DGX agent

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

DGX agent

arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Exist

safetyarxiv-cs-cl
27 Jul 2026
Safety

CommandLM: Data driven behavior level descriptor for ego vehicles

DGX agent

arXiv:2607.22078v1 Announce Type: new Abstract: As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decision-making is essential for safety, trust, and regul

safetyarxiv-cs-cv
27 Jul 2026
Safety

but they won’t.

DGX agent

but they won’t. AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better

safetygary-marcus--x
25 Jul 2026
Safety

End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

DGX agent

arXiv:2607.20674v1 Announce Type: new Abstract: We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier function

safetyarxiv-cs-lg
24 Jul 2026
Safety

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

DGX agent

arXiv:2607.07601v1 Announce Type: cross Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize co

safetyarxiv-cs-ai
9 Jul 2026
Safety

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

DGX agent

arXiv:2607.02569v1 Announce Type: new Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using

safetyarxiv-cs-cv
7 Jul 2026
Safety

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

DGX agent

arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained fil

safetyarxiv-cs-cl
3 Jul 2026
Safety

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

DGX agent

arXiv:2606.23671v2 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and e

safetyarxiv-cs-cl
30 Jun 2026
Safety

In-Vehicle Digital Twin-Based Collision Warning Framework with Sybil Attack Detection

DGX agent

arXiv:2606.28625v1 Announce Type: cross Abstract: Connected Vehicles (CVs) rely extensively on communication technologies to enable data-driven predictive analyses for enhancing performance and safety

safetyarxiv-cs-lg
30 Jun 2026
Safety

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

DGX agent

arXiv:2606.12342v1 Announce Type: cross Abstract: Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language.

safetyarxiv-cs-ai
11 Jun 2026
Safety

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment

DGX agent

arXiv:2510.03520v2 Announce Type: replace-cross Abstract: Ensuring safety is a foundational requirement for large language models (LLMs). Achieving an appropriate balance between enhancing the utility

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

DGX agent

arXiv:2606.05660v1 Announce Type: new Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety t

model-releasesarxiv-cs-ro
5 Jun 2026
Model Releases

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format

DGX agent

arXiv:2509.02655v3 Announce Type: replace-cross Abstract: Many AI alignment discussions of 'runaway optimisation' focus on RL agents: unbounded utility maximisers that over-optimise a proxy objective

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Constitutional On-Policy Safe Distillation

DGX agent

arXiv:2606.03089v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to prov

safetyarxiv-cs-ai
3 Jun 2026
Safety

COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

DGX agent

arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf

safetyarxiv-cs-ai
1 Jun 2026
Safety

Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment

DGX agent

arXiv:2605.27659v1 Announce Type: cross Abstract: Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicle

safetyarxiv-cs-ai
28 May 2026
Safety

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty

DGX agent

arXiv:2605.26627v1 Announce Type: cross Abstract: Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising

safetyarxiv-cs-ro
27 May 2026
← Previous
1…1617181920…297
Next →