AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Safety

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

DGX agent

arXiv:2606.05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow

safetyarxiv-cs-ai
6 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Where Should Knowledge Enter? A Layered Framework for Knowledge Infusion in Multimodal Iterative Generative Mo

DGX agent

arXiv:2606.06356v1 Announce Type: new Abstract: Multimodal generative models produce fluent outputs but remain unreliable when generation must respect structured, domain-specific, or safety-critical k

safetyarxiv-cs-ai
6 Jun 2026
Safety

Drishti AI-Event Guardian: An Intelligent Real-Time Crowd Monitoring and Emergency Response System for Mass Gathering Events

DGX agent

arXiv:2606.05185v1 Announce Type: cross Abstract: Mass gathering events are associated with critical safety incidents caused by insufficient crowd monitoring and inadequate emergency response coordina

safetyarxiv-cs-cv
5 Jun 2026
Safety

Instance-Level Post Hoc Uncertainty Quantification in Object Detection

DGX agent

arXiv:2606.04656v1 Announce Type: cross Abstract: Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for saf

safetyarxiv-cs-ai
4 Jun 2026
Safety

Multi-Agent Next-Best-View Optimization for Risk-Averse Planning

DGX agent

arXiv:2606.04158v1 Announce Type: new Abstract: Multi-agent Next-Best-View (NBV) selection for safe path planning in uncertain and unknown environments requires informative, safety-aware, and efficien

safetyarxiv-cs-ro
4 Jun 2026
Safety

Easy-to-Use Shielding for Reinforcement Learning

DGX agent

arXiv:2606.03804v1 Announce Type: new Abstract: Safe exploration is a key challenge in Reinforcement Learning (RL) that aims to prevent agents from making harmful decisions while exploring their envir

safetyarxiv-cs-lg
3 Jun 2026
Safety

Jailbreak Attack Initializations as Extractors of Compliance Directions

DGX agent

arXiv:2502.09755v4 Announce Type: replace-cross Abstract: Safety-aligned LLMs respond to prompts with either compliance or refusal, each corresponding to distinct directions in the model's activation

safetyarxiv-cs-lg
3 Jun 2026
Safety

Embedding Semantic Risk into Distance Fields and CBFs for Online Monocular Safe Control

DGX agent

arXiv:2606.01605v1 Announce Type: new Abstract: We propose an online monocular perception-to-control framework that embeds semantic risk into the distance field used by Control Barrier Function (CBF)-

safetyarxiv-cs-ro
2 Jun 2026
Safety

Multi-Objective Reinforcement Learning for Tactical Decision Making for Trucks in Highway Traffic

DGX agent

arXiv:2601.18783v2 Announce Type: replace-cross Abstract: Balancing safety, efficiency, and operational costs in highway driving poses a challenging decision-making problem for heavy-duty vehicles. A

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Predicted-Flow Control Barrier Functions for Real-Time Safe Optimal Control

DGX agent

arXiv:2606.00297v1 Announce Type: cross Abstract: Control barrier functions (CBFs) provide real-time safety guarantees through pointwise conditions on the state. However, synthesizing a valid CBF is d

model-releasesarxiv-cs-ro
2 Jun 2026
Safety

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

DGX agent

arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate

safetyarxiv-cs-cl
2 Jun 2026
Safety

Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

DGX agent

arXiv:2606.00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential a

safetyarxiv-cs-ai
2 Jun 2026
Safety

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer

DGX agent

arXiv:2602.02557v2 Announce Type: replace-cross Abstract: Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignmen

safetyarxiv-cs-ai
2 Jun 2026
Safety

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

DGX agent

arXiv:2602.07253v2 Announce Type: replace Abstract: Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existin

safetyarxiv-cs-ai
1 Jun 2026
Safety

Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems

DGX agent

This newsletter issue discusses three key topics in AI development and safety: the challenges involved in overseeing and controlling advanced AI systems, empirical findings about how protein folding A

safetyimport-ai
1 Jun 2026
Safety

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

DGX agent

arXiv:2512.11944v2 Announce Type: replace-cross Abstract: Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and int

safetyarxiv-cs-ai
29 May 2026
Safety

Evolving and Detecting Multi-Turn Deception using Geometric Signatures

DGX agent

arXiv:2605.27671v1 Announce Type: cross Abstract: Safety defenses for large language models (LLMs) are typically trained and evaluated on single-turn prompts, yet real attacks often unfold as indirect

safetyarxiv-cs-lg
28 May 2026
Safety

High-Fidelity Industrial Crash Dynamics Prediction via Geometry-Aware Operator Learning with Memory-Efficient Low-Rank Attention

DGX agent

arXiv:2605.27758v1 Announce Type: cross Abstract: Automotive crashworthiness optimization remains a safety-critical challenge, requiring the management of large-scale nonlinear structural deformations

safetyarxiv-cs-ai
28 May 2026
Safety

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

DGX agent

arXiv:2605.27157v1 Announce Type: new Abstract: Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robust

safetyarxiv-cs-ai
27 May 2026
Safety

DBPnet: Damper Characteristics-Based Bayesian Physics-Informed Neural Network for Wheel Load Estimation

DGX agent

arXiv:2605.24860v1 Announce Type: cross Abstract: Advanced driver assistance systems (ADAS) play an important role in modern automotive intelligence, significantly enhancing vehicle safety and stabili

safetyarxiv-cs-ai
26 May 2026
Safety

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

DGX agent

arXiv:2602.00491v2 Announce Type: replace Abstract: Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it r

safetyarxiv-cs-cl
26 May 2026
Safety

Empowering 9-1-1 Calltaking Training with Generative AI: Experiences and Lessons Learned

DGX agent

arXiv:2602.13241v2 Announce Type: replace-cross Abstract: Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a susta

safetyarxiv-cs-ai
25 May 2026
Safety

General Hazard Detection

DGX agent

arXiv:2605.23304v1 Announce Type: new Abstract: Hazard, as an abstract concept, is typically defined through cognitive-level logical reasoning rather than concrete examples. In contrast, existing haza

safetyarxiv-cs-cv
25 May 2026
Safety

Safe and Steerable Geometric Motion Policies for Robotic Dexterous Manipulation

DGX agent

arXiv:2605.21811v1 Announce Type: new Abstract: Robotic dexterous manipulation requires continuously reconciling objectives and constraints defined on heterogeneous geometric spaces: a robot controlle

safetyarxiv-cs-ro
22 May 2026
Safety

Reinforcement Learning for Risk Adaptation via Differentiable CVaR Barrier Functions

DGX agent

arXiv:2605.21257v1 Announce Type: new Abstract: Planning through crowded environments under uncertain obstacle motions remains difficult, as stochastic interactions often induce overly conservative be

safetyarxiv-cs-ro
21 May 2026
Safety

Exploring and Developing a Pre-Model Safeguard with Draft Models

DGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

safetyarxiv-cs-ai
20 May 2026
Safety

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails

DGX agent

arXiv:2510.13727v2 Announce Type: replace Abstract: Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-g

safetyarxiv-cs-ai
20 May 2026
Safety

Generative Auto-Bidding with Unified Modeling and Exploration

DGX agent

arXiv:2605.19457v1 Announce Type: new Abstract: Automated bidding is central to modern digital advertising. Early rule-based methods lacked adaptability, while subsequent Reinforcement Learning approa

safetyarxiv-cs-ai
20 May 2026
Safety

k-Inductive Neural Barrier Certificates for Unknown Nonlinear Dynamics

DGX agent

arXiv:2605.20108v1 Announce Type: cross Abstract: While conventional (k=1) discrete-time barrier certificate conditions impose strict safety constraints by requiring the function to be non-increasing

safetyarxiv-cs-ai
20 May 2026
Safety

AI Alignment Breaks at the Edge

DGX agent

arXiv:2602.20042v2 Announce Type: replace Abstract: General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. T

safetyarxiv-cs-cl
19 May 2026
Safety

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

DGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

safetyarxiv-cs-cl
19 May 2026
Safety

New Wide-Net-Casting Jailbreak Attacks Risk Large Models

DGX agent

arXiv:2605.17128v1 Announce Type: cross Abstract: Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexpl

safetyarxiv-cs-ai
19 May 2026
Safety

Bellman Value Decomposition for Task Logic in Safe Optimal Control

DGX agent

arXiv:2602.19532v2 Announce Type: replace Abstract: Real-world tasks involve nuanced combinations of goal and safety specifications. In high dimensions, the challenge is exacerbated: formal automata b

safetyarxiv-cs-ro
15 May 2026
Safety

GradShield: Alignment Preserving Finetuning

DGX agent

arXiv:2605.14194v1 Announce Type: new Abstract: Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implic

safetyarxiv-cs-cl
15 May 2026
Safety

Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study

DGX agent

arXiv:2605.14087v1 Announce Type: new Abstract: Large Language Models (LLMs), when trained on web-scale corpora, inherently absorb toxic patterns from their training data. This leads to ``toxic degene

safetyarxiv-cs-cl
15 May 2026
Safety

A Five-Layer MLOps Architecture for Connected Automated Driving

DGX agent

arXiv:2605.12719v1 Announce Type: cross Abstract: The continual assurance of safety and performance of automated driving systems (ADSs) poses significant challenges. ADSs operate in complex, dynamic,

safetyarxiv-cs-lg
14 May 2026
Safety

Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification

DGX agent

arXiv:2605.12917v1 Announce Type: new Abstract: Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Predicti

safetyarxiv-cs-cv
14 May 2026
Safety

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

DGX agent

arXiv:2605.13755v1 Announce Type: new Abstract: In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critic

safetyarxiv-cs-cv
14 May 2026
Safety

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

DGX agent

arXiv:2605.13801v1 Announce Type: cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of th

safetyarxiv-cs-ai
14 May 2026
Safety

MoCCA: A Movable Circle Probability of Collision Approximation

DGX agent

arXiv:2605.13125v1 Announce Type: new Abstract: In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and

safetyarxiv-cs-ro
14 May 2026
Safety

Tracing Persona Vectors Through LLM Pretraining

DGX agent

arXiv:2605.13329v1 Announce Type: cross Abstract: How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determi

safetyarxiv-cs-ai
14 May 2026
Safety

Cooperative Robotics Reinforced by Collective Perception for Traffic Moderation

DGX agent

arXiv:2605.11972v1 Announce Type: new Abstract: Collisions at non-line-of-sight (NLOS) intersections remain a major safety concern because drivers have limited visibility of approaching traffic. V2X b

safetyarxiv-cs-ro
13 May 2026
Safety

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

DGX agent

arXiv:2411.16769v3 Announce Type: replace-cross Abstract: Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, hum

safetyarxiv-cs-cl
13 May 2026
Safety

Robustness Certificates for Neural Networks against Adversarial Attacks

DGX agent

arXiv:2512.20865v2 Announce Type: replace Abstract: The increasing use of machine learning in safety-critical domains amplifies the risk of adversarial threats, especially data poisoning attacks that

safetyarxiv-cs-lg
13 May 2026
Safety

An Empirical Analysis of Calibration and Selective Prediction in Multimodal Clinical Condition Classification

DGX agent

arXiv:2603.02719v2 Announce Type: replace Abstract: As artificial intelligence systems move toward clinical deployment, ensuring reliable prediction behavior is fundamental for safety-critical decisio

safetyarxiv-cs-lg
12 May 2026
Safety

Conformity Generates Collective Misalignment in AI Agents Societies

DGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

safetyarxiv-cs-cl
12 May 2026
Safety

Hierarchical Causal Abduction: A Foundation Framework for Explainable Model Predictive Control

DGX agent

arXiv:2605.10624v1 Announce Type: new Abstract: Model Predictive Control (MPC) is widely used to operate safety-critical infrastructure by predicting future trajectories and optimizing control actions

safetyarxiv-cs-ai
12 May 2026
Safety

Learning When to Jump for Off-road Navigation

DGX agent

arXiv:2602.00877v2 Announce Type: replace Abstract: Low speed does not always guarantee safety in off-road driving. For instance, crossing a ditch may be risky at a low speed due to the risk of gettin

safetyarxiv-cs-ro
12 May 2026
← Previous
1…2122232425…299
Next →