AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

General Hazard Detection

DGX agent

arXiv:2605.23304v1 Announce Type: new Abstract: Hazard, as an abstract concept, is typically defined through cognitive-level logical reasoning rather than concrete examples. In contrast, existing haza

safetyarxiv-cs-cv
25 May 2026
Safety

Safe and Steerable Geometric Motion Policies for Robotic Dexterous Manipulation

DGX agent

arXiv:2605.21811v1 Announce Type: new Abstract: Robotic dexterous manipulation requires continuously reconciling objectives and constraints defined on heterogeneous geometric spaces: a robot controlle

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
safetyarxiv-cs-ro
22 May 2026
Safety

Reinforcement Learning for Risk Adaptation via Differentiable CVaR Barrier Functions

DGX agent

arXiv:2605.21257v1 Announce Type: new Abstract: Planning through crowded environments under uncertain obstacle motions remains difficult, as stochastic interactions often induce overly conservative be

safetyarxiv-cs-ro
21 May 2026
Safety

Exploring and Developing a Pre-Model Safeguard with Draft Models

DGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

safetyarxiv-cs-ai
20 May 2026
Safety

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails

DGX agent

arXiv:2510.13727v2 Announce Type: replace Abstract: Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-g

safetyarxiv-cs-ai
20 May 2026
Safety

Generative Auto-Bidding with Unified Modeling and Exploration

DGX agent

arXiv:2605.19457v1 Announce Type: new Abstract: Automated bidding is central to modern digital advertising. Early rule-based methods lacked adaptability, while subsequent Reinforcement Learning approa

safetyarxiv-cs-ai
20 May 2026
Safety

k-Inductive Neural Barrier Certificates for Unknown Nonlinear Dynamics

DGX agent

arXiv:2605.20108v1 Announce Type: cross Abstract: While conventional (k=1) discrete-time barrier certificate conditions impose strict safety constraints by requiring the function to be non-increasing

safetyarxiv-cs-ai
20 May 2026
Safety

AI Alignment Breaks at the Edge

DGX agent

arXiv:2602.20042v2 Announce Type: replace Abstract: General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. T

safetyarxiv-cs-cl
19 May 2026
Safety

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

DGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

safetyarxiv-cs-cl
19 May 2026
Safety

New Wide-Net-Casting Jailbreak Attacks Risk Large Models

DGX agent

arXiv:2605.17128v1 Announce Type: cross Abstract: Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexpl

safetyarxiv-cs-ai
19 May 2026
Safety

Bellman Value Decomposition for Task Logic in Safe Optimal Control

DGX agent

arXiv:2602.19532v2 Announce Type: replace Abstract: Real-world tasks involve nuanced combinations of goal and safety specifications. In high dimensions, the challenge is exacerbated: formal automata b

safetyarxiv-cs-ro
15 May 2026
Safety

GradShield: Alignment Preserving Finetuning

DGX agent

arXiv:2605.14194v1 Announce Type: new Abstract: Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implic

safetyarxiv-cs-cl
15 May 2026
Safety

Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study

DGX agent

arXiv:2605.14087v1 Announce Type: new Abstract: Large Language Models (LLMs), when trained on web-scale corpora, inherently absorb toxic patterns from their training data. This leads to ``toxic degene

safetyarxiv-cs-cl
15 May 2026
Safety

A Five-Layer MLOps Architecture for Connected Automated Driving

DGX agent

arXiv:2605.12719v1 Announce Type: cross Abstract: The continual assurance of safety and performance of automated driving systems (ADSs) poses significant challenges. ADSs operate in complex, dynamic,

safetyarxiv-cs-lg
14 May 2026
Safety

Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification

DGX agent

arXiv:2605.12917v1 Announce Type: new Abstract: Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Predicti

safetyarxiv-cs-cv
14 May 2026
Safety

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

DGX agent

arXiv:2605.13755v1 Announce Type: new Abstract: In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critic

safetyarxiv-cs-cv
14 May 2026
Safety

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

DGX agent

arXiv:2605.13801v1 Announce Type: cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of th

safetyarxiv-cs-ai
14 May 2026
Safety

MoCCA: A Movable Circle Probability of Collision Approximation

DGX agent

arXiv:2605.13125v1 Announce Type: new Abstract: In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and

safetyarxiv-cs-ro
14 May 2026
Safety

Tracing Persona Vectors Through LLM Pretraining

DGX agent

arXiv:2605.13329v1 Announce Type: cross Abstract: How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determi

safetyarxiv-cs-ai
14 May 2026
Safety

Cooperative Robotics Reinforced by Collective Perception for Traffic Moderation

DGX agent

arXiv:2605.11972v1 Announce Type: new Abstract: Collisions at non-line-of-sight (NLOS) intersections remain a major safety concern because drivers have limited visibility of approaching traffic. V2X b

safetyarxiv-cs-ro
13 May 2026
Safety

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

DGX agent

arXiv:2411.16769v3 Announce Type: replace-cross Abstract: Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, hum

safetyarxiv-cs-cl
13 May 2026
Safety

Robustness Certificates for Neural Networks against Adversarial Attacks

DGX agent

arXiv:2512.20865v2 Announce Type: replace Abstract: The increasing use of machine learning in safety-critical domains amplifies the risk of adversarial threats, especially data poisoning attacks that

safetyarxiv-cs-lg
13 May 2026
Safety

An Empirical Analysis of Calibration and Selective Prediction in Multimodal Clinical Condition Classification

DGX agent

arXiv:2603.02719v2 Announce Type: replace Abstract: As artificial intelligence systems move toward clinical deployment, ensuring reliable prediction behavior is fundamental for safety-critical decisio

safetyarxiv-cs-lg
12 May 2026
Safety

Conformity Generates Collective Misalignment in AI Agents Societies

DGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

safetyarxiv-cs-cl
12 May 2026
Safety

Hierarchical Causal Abduction: A Foundation Framework for Explainable Model Predictive Control

DGX agent

arXiv:2605.10624v1 Announce Type: new Abstract: Model Predictive Control (MPC) is widely used to operate safety-critical infrastructure by predicting future trajectories and optimizing control actions

safetyarxiv-cs-ai
12 May 2026
Safety

Learning When to Jump for Off-road Navigation

DGX agent

arXiv:2602.00877v2 Announce Type: replace Abstract: Low speed does not always guarantee safety in off-road driving. For instance, crossing a ditch may be risky at a low speed due to the risk of gettin

safetyarxiv-cs-ro
12 May 2026
Safety

Mismatch-Aware Adaptive Constraint Tightening for Bicycle-Model Trajectory Optimization

DGX agent

arXiv:2605.09376v1 Announce Type: new Abstract: Trajectory optimization for autonomous vehicles usually relies on the kinematic bicycle model because of its computational simplicity. However, when the

safetyarxiv-cs-ro
12 May 2026
Safety

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

DGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

safetyarxiv-cs-ai
12 May 2026
Safety

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

DGX agent

arXiv:2605.07324v1 Announce Type: cross Abstract: Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior w

safetyarxiv-cs-ai
11 May 2026
Safety

A Vision-Based Shared-Control Teleoperation Scheme for Controlling the Robotic Arm of a Four-Legged Robot

DGX agent

arXiv:2508.14994v3 Announce Type: replace-cross Abstract: In hazardous and remote environments, robotic systems perform critical tasks demanding improved safety and efficiency. Among these, quadruped

safetyarxiv-cs-cv
6 May 2026
Safety

EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models

DGX agent

arXiv:2605.02921v1 Announce Type: cross Abstract: As LLMs continue to shape real-world applications, automated jailbreak generation becomes essential to reveal safety weaknesses and guide model improv

safetyarxiv-cs-lg
6 May 2026
Safety

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

DGX agent

arXiv:2601.07885v2 Announce Type: replace-cross Abstract: Emoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs)

safetyarxiv-cs-ai
6 May 2026
Safety

Autonomous Reliability Qualification of Ga_2O_3-based Hydrogen and Temperature Sensors via Safe Active Learning

DGX agent

arXiv:2605.00868v1 Announce Type: cross Abstract: We present a Safe Active Learning (SAL) framework for autonomous reliability characterization of rectifying Ga_2O_3-based devices under coupled therma

safetyarxiv-cs-lg
5 May 2026
Safety

Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit

DGX agent

arXiv:2602.07249v2 Announce Type: replace-cross Abstract: Autonomous Vehicles (AVs), especially vision-based AVs, are rapidly being deployed without human operators. As AVs operate in safety-critical

safetyarxiv-cs-lg
5 May 2026
Safety

Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation

DGX agent

arXiv:2605.01113v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models have the ability to build high-quality pictures from text prompts, but they pose safety concerns because they can g

safetyarxiv-cs-cv
5 May 2026
Safety

Evidence-Based Landing Site Selection and Vison-Based Landing for UAVs in Unstructured Environments

DGX agent

arXiv:2605.01432v1 Announce Type: new Abstract: Autonomous landing in cluttered or unstructured environments remains a safety-critical challenge for unmanned aerial vehicles (UAVs), particularly under

safetyarxiv-cs-ro
5 May 2026
Safety

A Deep Learning-Based CCTV System for Automatic Smoking Detection in Fire Exit Zones

DGX agent

arXiv:2508.11696v3 Announce Type: replace Abstract: A deep learning real-time smoking detection system for CCTV surveillance of fire exit areas is proposed due to critical safety requirements. The dat

safetyarxiv-cs-cv
4 May 2026
Safety

Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models

DGX agent

arXiv:2602.03265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms.

safetyarxiv-cs-lg
4 May 2026
Safety

OpenAI o1 System Card

DGX agent

arXiv:2412.16720v2 Announce Type: replace Abstract: The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provi

safetyarxiv-cs-ai
1 May 2026
Safety

No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control

DGX agent

arXiv:2604.25887v1 Announce Type: new Abstract: Current pedestrian crossing signals operate on fixed timing without adjustment to pedestrian behavior, which can leave vulnerable road users (VRUs) such

safetyarxiv-cs-cv
29 Apr 2026
Safety

A Self-Supervised Framework for Space Object Behaviour Characterisation

DGX agent

arXiv:2504.06176v3 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly

safetyarxiv-cs-ai
28 Apr 2026
Safety

FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data

DGX agent

arXiv:2604.24572v1 Announce Type: new Abstract: The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM), maintained by the Observational Health Data Sciences and Informatics (OHDSI

safetyarxiv-cs-ai
28 Apr 2026
Safety

Pedestrians play chicken with an autonomous vehicle

DGX agent

arXiv:2604.24384v1 Announce Type: new Abstract: Automated vehicles (AVs) are commonly programmed to yield unconditionally to pedestrians in the interest of safety. However, this design choice can give

safetyarxiv-cs-ro
28 Apr 2026
Safety

Risk-Aware Robust Learning: Reducing Clinical Risk under Label Noise in Medical Image Classification

DGX agent

arXiv:2604.23875v1 Announce Type: cross Abstract: Noisy labels are a pervasive challenge in medical image classification, where annotation errors arise from inter-observer variability and diagnostic a

safetyarxiv-cs-ai
28 Apr 2026
Safety

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

DGX agent

arXiv:2604.22103v1 Announce Type: cross Abstract: Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized vi

safetyarxiv-cs-cv
27 Apr 2026
Safety

Survey on Evaluation of LLM-based Agents

DGX agent

arXiv:2503.16416v2 Announce Type: replace Abstract: LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environ

safetyarxiv-cs-ai
24 Apr 2026
Safety

Aligning Human-AI-Interaction Trust for Mental Health Support: Survey and Position for Multi-Stakeholders

DGX agent

arXiv:2604.20166v1 Announce Type: new Abstract: Building trustworthy AI systems for mental health support is a shared priority across stakeholders from multiple disciplines. However, 'trustworthy' rem

safetyarxiv-cs-cl
23 Apr 2026
Safety

Interval POMDP Shielding for Imperfect-Perception Agents

DGX agent

arXiv:2604.20728v1 Announce Type: new Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for this setting

safetyarxiv-cs-ai
23 Apr 2026
← Previous
1…1920212223…257
Next →