AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
21 May 2026

Reinforcement Learning for Risk Adaptation via Differentiable CVaR Barrier Functions

SafetyDGX agent

arXiv:2605.21257v1 Announce Type: new Abstract: Planning through crowded environments under uncertain obstacle motions remains difficult, as stochastic interactions often induce overly conservative be

20 May 2026

Exploring and Developing a Pre-Model Safeguard with Draft Models

SafetyDGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails

Safety
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2510.13727v2 Announce Type: replace Abstract: Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-g

Generative Auto-Bidding with Unified Modeling and Exploration

SafetyDGX agent

arXiv:2605.19457v1 Announce Type: new Abstract: Automated bidding is central to modern digital advertising. Early rule-based methods lacked adaptability, while subsequent Reinforcement Learning approa

k-Inductive Neural Barrier Certificates for Unknown Nonlinear Dynamics

SafetyDGX agent

arXiv:2605.20108v1 Announce Type: cross Abstract: While conventional (k=1) discrete-time barrier certificate conditions impose strict safety constraints by requiring the function to be non-increasing

19 May 2026

AI Alignment Breaks at the Edge

SafetyDGX agent

arXiv:2602.20042v2 Announce Type: replace Abstract: General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. T

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

SafetyDGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

New Wide-Net-Casting Jailbreak Attacks Risk Large Models

SafetyDGX agent

arXiv:2605.17128v1 Announce Type: cross Abstract: Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexpl

15 May 2026

Bellman Value Decomposition for Task Logic in Safe Optimal Control

SafetyDGX agent

arXiv:2602.19532v2 Announce Type: replace Abstract: Real-world tasks involve nuanced combinations of goal and safety specifications. In high dimensions, the challenge is exacerbated: formal automata b

GradShield: Alignment Preserving Finetuning

SafetyDGX agent

arXiv:2605.14194v1 Announce Type: new Abstract: Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implic

Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study

SafetyDGX agent

arXiv:2605.14087v1 Announce Type: new Abstract: Large Language Models (LLMs), when trained on web-scale corpora, inherently absorb toxic patterns from their training data. This leads to ``toxic degene

14 May 2026

A Five-Layer MLOps Architecture for Connected Automated Driving

SafetyDGX agent

arXiv:2605.12719v1 Announce Type: cross Abstract: The continual assurance of safety and performance of automated driving systems (ADSs) poses significant challenges. ADSs operate in complex, dynamic,

Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification

SafetyDGX agent

arXiv:2605.12917v1 Announce Type: new Abstract: Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Predicti

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

SafetyDGX agent

arXiv:2605.13755v1 Announce Type: new Abstract: In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critic

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

SafetyDGX agent

arXiv:2605.13801v1 Announce Type: cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of th

MoCCA: A Movable Circle Probability of Collision Approximation

SafetyDGX agent

arXiv:2605.13125v1 Announce Type: new Abstract: In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and

Tracing Persona Vectors Through LLM Pretraining

SafetyDGX agent

arXiv:2605.13329v1 Announce Type: cross Abstract: How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determi

13 May 2026

Cooperative Robotics Reinforced by Collective Perception for Traffic Moderation

SafetyDGX agent

arXiv:2605.11972v1 Announce Type: new Abstract: Collisions at non-line-of-sight (NLOS) intersections remain a major safety concern because drivers have limited visibility of approaching traffic. V2X b

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

SafetyDGX agent

arXiv:2411.16769v3 Announce Type: replace-cross Abstract: Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, hum

Robustness Certificates for Neural Networks against Adversarial Attacks

SafetyDGX agent

arXiv:2512.20865v2 Announce Type: replace Abstract: The increasing use of machine learning in safety-critical domains amplifies the risk of adversarial threats, especially data poisoning attacks that

12 May 2026

An Empirical Analysis of Calibration and Selective Prediction in Multimodal Clinical Condition Classification

SafetyDGX agent

arXiv:2603.02719v2 Announce Type: replace Abstract: As artificial intelligence systems move toward clinical deployment, ensuring reliable prediction behavior is fundamental for safety-critical decisio

Conformity Generates Collective Misalignment in AI Agents Societies

SafetyDGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

Hierarchical Causal Abduction: A Foundation Framework for Explainable Model Predictive Control

SafetyDGX agent

arXiv:2605.10624v1 Announce Type: new Abstract: Model Predictive Control (MPC) is widely used to operate safety-critical infrastructure by predicting future trajectories and optimizing control actions

Learning When to Jump for Off-road Navigation

SafetyDGX agent

arXiv:2602.00877v2 Announce Type: replace Abstract: Low speed does not always guarantee safety in off-road driving. For instance, crossing a ditch may be risky at a low speed due to the risk of gettin

Mismatch-Aware Adaptive Constraint Tightening for Bicycle-Model Trajectory Optimization

SafetyDGX agent

arXiv:2605.09376v1 Announce Type: new Abstract: Trajectory optimization for autonomous vehicles usually relies on the kinematic bicycle model because of its computational simplicity. However, when the

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

SafetyDGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

Sam Altman swearing to tell the whole truth, and then failing to do so. May 2023.

SafetyDGX agent

Sam Altman made statements under oath in May 2023 regarding AI safety and OpenAI's practices, but Gary Marcus critiqued these statements as incomplete or misleading, suggesting Altman failed to fully

11 May 2026

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

SafetyDGX agent

arXiv:2605.07324v1 Announce Type: cross Abstract: Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior w

8 May 2026

Good summary of today, @katiemiller, but then again it is getting hard to track of the total number of ex-board members who have called Altm…

SafetyDGX agent

Good summary of today, @katiemiller, but then again it is getting hard to track of the total number of ex-board members who have called Altman a liar 🤷‍♂️ Also hard to keep track how many OpenAI safet

6 May 2026

A Vision-Based Shared-Control Teleoperation Scheme for Controlling the Robotic Arm of a Four-Legged Robot

SafetyDGX agent

arXiv:2508.14994v3 Announce Type: replace-cross Abstract: In hazardous and remote environments, robotic systems perform critical tasks demanding improved safety and efficiency. Among these, quadruped

EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models

SafetyDGX agent

arXiv:2605.02921v1 Announce Type: cross Abstract: As LLMs continue to shape real-world applications, automated jailbreak generation becomes essential to reveal safety weaknesses and guide model improv

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

SafetyDGX agent

arXiv:2601.07885v2 Announce Type: replace-cross Abstract: Emoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs)

5 May 2026

Autonomous Reliability Qualification of Ga_2O_3-based Hydrogen and Temperature Sensors via Safe Active Learning

SafetyDGX agent

arXiv:2605.00868v1 Announce Type: cross Abstract: We present a Safe Active Learning (SAL) framework for autonomous reliability characterization of rectifying Ga_2O_3-based devices under coupled therma

Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit

SafetyDGX agent

arXiv:2602.07249v2 Announce Type: replace-cross Abstract: Autonomous Vehicles (AVs), especially vision-based AVs, are rapidly being deployed without human operators. As AVs operate in safety-critical

Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation

SafetyDGX agent

arXiv:2605.01113v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models have the ability to build high-quality pictures from text prompts, but they pose safety concerns because they can g

Evidence-Based Landing Site Selection and Vison-Based Landing for UAVs in Unstructured Environments

SafetyDGX agent

arXiv:2605.01432v1 Announce Type: new Abstract: Autonomous landing in cluttered or unstructured environments remains a safety-critical challenge for unmanned aerial vehicles (UAVs), particularly under

4 May 2026

A Deep Learning-Based CCTV System for Automatic Smoking Detection in Fire Exit Zones

SafetyDGX agent

arXiv:2508.11696v3 Announce Type: replace Abstract: A deep learning real-time smoking detection system for CCTV surveillance of fire exit areas is proposed due to critical safety requirements. The dat

Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models

SafetyDGX agent

arXiv:2602.03265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms.

1 May 2026

OpenAI o1 System Card

SafetyDGX agent

arXiv:2412.16720v2 Announce Type: replace Abstract: The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provi

29 Apr 2026

No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control

SafetyDGX agent

arXiv:2604.25887v1 Announce Type: new Abstract: Current pedestrian crossing signals operate on fixed timing without adjustment to pedestrian behavior, which can leave vulnerable road users (VRUs) such

28 Apr 2026

A Self-Supervised Framework for Space Object Behaviour Characterisation

SafetyDGX agent

arXiv:2504.06176v3 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly

FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data

SafetyDGX agent

arXiv:2604.24572v1 Announce Type: new Abstract: The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM), maintained by the Observational Health Data Sciences and Informatics (OHDSI

Pedestrians play chicken with an autonomous vehicle

SafetyDGX agent

arXiv:2604.24384v1 Announce Type: new Abstract: Automated vehicles (AVs) are commonly programmed to yield unconditionally to pedestrians in the interest of safety. However, this design choice can give

Risk-Aware Robust Learning: Reducing Clinical Risk under Label Noise in Medical Image Classification

SafetyDGX agent

arXiv:2604.23875v1 Announce Type: cross Abstract: Noisy labels are a pervasive challenge in medical image classification, where annotation errors arise from inter-observer variability and diagnostic a

27 Apr 2026

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

SafetyDGX agent

arXiv:2604.22103v1 Announce Type: cross Abstract: Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized vi

If extremely violent criminals are not imprisoned, eventually they will murder innocent people

SafetyDGX agent

If extremely violent criminals are not imprisoned, eventually they will murder innocent people This bodega owner told ABC a year ago that he fears for his safety in NY Last night, he was kiIIed by a s

24 Apr 2026

Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the d…

SafetyDGX agent

Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the delivery of information and entertainment and the frictionles

Survey on Evaluation of LLM-based Agents

SafetyDGX agent

arXiv:2503.16416v2 Announce Type: replace Abstract: LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environ

23 Apr 2026

Aligning Human-AI-Interaction Trust for Mental Health Support: Survey and Position for Multi-Stakeholders

SafetyDGX agent

arXiv:2604.20166v1 Announce Type: new Abstract: Building trustworthy AI systems for mental health support is a shared priority across stakeholders from multiple disciplines. However, 'trustworthy' rem

Interval POMDP Shielding for Imperfect-Perception Agents

SafetyDGX agent

arXiv:2604.20728v1 Announce Type: new Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for this setting

22 Apr 2026

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

SafetyDGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

SafetyDGX agent

arXiv:2510.26782v3 Announce Type: replace-cross Abstract: A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical

Integrating Anomaly Detection into Agentic AI for Proactive Risk Management in Human Activity

SafetyDGX agent

arXiv:2604.19538v1 Announce Type: new Abstract: Agentic AI, with goal-directed, proactive, and autonomous decision-making capabilities, offers a compelling opportunity to address movement-related risk

Large Language Models Exhibit Normative Conformity

SafetyDGX agent

arXiv:2604.19301v1 Announce Type: new Abstract: The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-

21 Apr 2026

Jailbreaking Large Language Models with Morality Attacks

SafetyDGX agent

arXiv:2604.17053v1 Announce Type: new Abstract: Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Resea

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

SafetyDGX agent

arXiv:2604.16358v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent through the evolving visual-text history and exploi

STL-Based Motion Planning and Uncertainty-Aware Risk Analysis for Human-Robot Collaboration with a Multi-Rotor Aerial Vehicle

SafetyDGX agent

arXiv:2509.10692v3 Announce Type: replace Abstract: This paper presents a motion planning and risk analysis framework for enhancing human-robot collaboration with a Multi-Rotor Aerial Vehicle. The pro

20 Apr 2026

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning

SafetyDGX agent

arXiv:2604.15705v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with comple

17 Apr 2026

SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models

SafetyDGX agent

arXiv:2604.14163v1 Announce Type: new Abstract: Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergencies at sea.

16 Apr 2026

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Te…

SafetyDGX agent

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Tesla-driven robotaxis. • Waymo: 58 incidents • Zoox: 3 incide

← Previous
1…1718192021…240
Next →