AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
13 May 2026

Prototype Fusion: A Training-Free Multi-Layer Approach to OOD Detection

SafetyDGX agent

arXiv:2603.23677v2 Announce Type: replace Abstract: Deep learning models are increasingly deployed in safety-critical applications, where reliable out-of-distribution (OOD) detection is essential to e

The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested

SafetyDGX agent

arXiv:2605.11496v1 Announce Type: cross Abstract: Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and

Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness

SafetyDGX agent

arXiv:2503.16072v4 Announce Type: replace-cross Abstract: Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet mo

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VNDUQE: Information-Theoretic Novelty Detection using Deep Variational Information Bottleneck

SafetyDGX agent

arXiv:2605.11551v1 Announce Type: cross Abstract: Detecting out-of-distribution (OOD) samples is critical for safe deployment of neural networks in safety-critical applications. While maximum softmax

12 May 2026

Agent-Sentry: Bounding LLM Agents via Execution Provenance

SafetyDGX agent

arXiv:2603.22868v2 Announce Type: replace-cross Abstract: Agentic computing systems, while immensely capable, raise serious security, privacy, and safety concerns. A key issue is that the full set of

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

SafetyDGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

Consensus Sampling for Safer Generative AI

SafetyDGX agent

arXiv:2511.09493v2 Announce Type: replace Abstract: Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions

Constraint-Aware Reinforcement Learning via Adaptive Action Scaling

SafetyDGX agent

arXiv:2510.11491v3 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) seeks to mitigate unsafe behaviors that arise from exploration during training by reducing constraint violati

GNN for Structural Displacement Prediction

SafetyDGX agent

arXiv:2605.08303v1 Announce Type: cross Abstract: Accurate prediction of structural displacements under external loading is fundamental to structural health monitoring and seismic safety assessment. A

How LLMs Are Persuaded: A Few Attention Heads, Rerouted

SafetyDGX agent

arXiv:2605.09314v1 Announce Type: new Abstract: Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly und

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

Model ReleasesDGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

SafetyDGX agent

arXiv:2602.03677v2 Announce Type: replace Abstract: Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliab

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

Model ReleasesDGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

SafetyDGX agent

arXiv:2605.05682v2 Announce Type: replace-cross Abstract: Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI

Positive Alignment: Artificial Intelligence for Human Flourishing

SafetyDGX agent

arXiv:2605.10310v1 Announce Type: new Abstract: Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of ali

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

SafetyDGX agent

arXiv:2605.10293v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide gua

Safe Exploration for Nonlinear Processes Using Online Gaussian Process Learning

SafetyDGX agent

arXiv:2605.09772v1 Announce Type: cross Abstract: This paper proposes a safe data-driven control framework for nonlinear systems with partially known dynamics. The method ensures stability and constra

UAV-Assisted Scan-to-Simulation for Landslides Using Physics-Informed Gaussian Splatting

SafetyDGX agent

arXiv:2605.10715v1 Announce Type: new Abstract: Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines t

VISTA: A Generative Egocentric Video Framework for Daily Assistance

SafetyDGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

SafetyDGX agent

arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne

11 May 2026

A Systematic Investigation of The RL-Jailbreaker in LLMs

SafetyDGX agent

arXiv:2605.07032v1 Announce Type: cross Abstract: The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversa

Approximation-Free Differentiable Oblique Decision Trees

SafetyDGX agent

arXiv:2605.07837v1 Announce Type: cross Abstract: Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabu

BEAVER: An Efficient Deterministic LLM Verifier

SafetyDGX agent

arXiv:2512.05439v2 Announce Type: replace Abstract: As large language models (LLMs) transition from research prototypes to production systems, practitioners often need reliable methods to verify model

Beyond 'I cannot fulfill this request': Alleviating Rigid Rejection in LLMs via Label Enhancement

SafetyDGX agent

arXiv:2605.07883v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often l

Fortifying Time Series: DTW-Certified Robust Anomaly Detection

SafetyDGX agent

arXiv:2605.07690v1 Announce Type: new Abstract: Time-series anomaly detection is critical for ensuring safety in high-stakes applications, where robustness is a fundamental requirement rather than a m

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines

SafetyDGX agent

arXiv:2605.07062v1 Announce Type: cross Abstract: AI agents are assuming active roles in Continuous Integration and Continuous Deployment (CI/CD) workflows, yet the research community lacks a shared v

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

SafetyDGX agent

arXiv:2605.06696v1 Announce Type: new Abstract: Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. Howev

How Value Induction Reshapes LLM Behaviour

SafetyDGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

Many-to-Many Multi-Agent Pickup and Delivery

SafetyDGX agent

arXiv:2605.07835v1 Announce Type: new Abstract: Multi-robot systems in automated warehouses must manage continuous streams of pickup-and-delivery tasks while ensuring efficiency and safety. Prior work

PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection

SafetyDGX agent

arXiv:2509.26272v3 Announce Type: replace Abstract: The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the

10 May 2026

So many bots but honestly decent alignment

SafetyDGX agent

This post likely discusses the proliferation of AI bots and chatbot applications while arguing that despite their abundance, many demonstrate reasonable safety practices or alignment with human values

7 May 2026

From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation

SafetyDGX agent

arXiv:2605.04327v1 Announce Type: new Abstract: We propose an architecture for integrating high-level, human-provided safety rules and operator-aligned semantic preferences into autonomous robot navig

Hey @bloomberg could you update this? It’s already way out of date. 🙏

SafetyDGX agent

Gary Marcus posted on X (formerly Twitter) requesting that Bloomberg update outdated information, likely regarding AI safety, regulation, or related technological topics that Marcus frequently discuss

LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy

SafetyDGX agent

arXiv:2605.04295v1 Announce Type: new Abstract: LLMs' overconfidence, particularly when hallucinating, poses a significant challenge for the deployment of the models in safety-critical settings and ma

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate

SafetyDGX agent

arXiv:2605.03409v1 Announce Type: new Abstract: We present Robust Agent Compensation (RAC), a log-based recovery paradigm (providing a safety net) implemented through an architectural extension that c

6 May 2026

A Three-Stage Offline SDRE-Based Control Framework for Human Motion Reproduction on a Suspended Bipedal Robot

SafetyDGX agent

arXiv:2506.04680v2 Announce Type: replace Abstract: During the development of wearable exoskeletons, evaluations involving human subjects pose inherent safety risks. Therefore, systematic testing is o

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

SafetyDGX agent

arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing

Efficient Temporal Datalog Materialisation for Composite Event Recognition

SafetyDGX agent

arXiv:2605.02488v1 Announce Type: new Abstract: Several applications demand the timely detection of critical situations, such as threats to safety and transparency, over high-velocity streams of symbo

Human-in-the-Loop Uncertainty Analysis in Self-Adaptive Robots Using LLMs

SafetyDGX agent

arXiv:2605.02983v1 Announce Type: new Abstract: Self-adaptive robots operate in dynamic, unpredictable environments where unaddressed uncertainties can lead to safety violations and operational failur

Instance-Level Costs for Nuanced Classifier Evaluation

SafetyDGX agent

arXiv:2605.03135v1 Announce Type: new Abstract: Standard classification treats all errors equally, but in content moderation, medical screening, and safety-critical applications, mistakes on clear-cut

Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings

SafetyDGX agent

arXiv:2605.02908v1 Announce Type: new Abstract: Understanding how textual embeddings contribute to memorization in text-to-image diffusion models is crucial for both interpretability and safety. This

Mira Murati tells the court that she couldn’t trust Sam Altman’s words

SafetyDGX agent

Mira Murati, OpenAI's former CTO, has testified under oath that CEO Sam Altman lied to her about the safety standards for a new AI model. In a video deposition shown during the ongoing Musk v. Altman

Mira Murati’s testimony is gripping – and what it makes absolutely clear is how utterly wrong most of Twitter was about why Sam was fired. –…

SafetyDGX agent

Mira Murati’s testimony is gripping – and what it makes absolutely clear is how utterly wrong most of Twitter was about why Sam was fired. – It had nothing per se to do with AI safety - It had nothing

Model Routing as a Trust Problem: Route Receipts for Adaptive AI Systems

SafetyDGX agent

arXiv:2605.01710v1 Announce Type: new Abstract: AI products often route requests through version aliases, service tiers, tool choices, regional endpoints, fallback rules, or safety handling before res

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs

Local AiDGX agent

arXiv:2605.02946v1 Announce Type: new Abstract: Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly

Steerable Adversarial Scenario Generation through Test-Time Preference Alignment

SafetyDGX agent

arXiv:2509.20102v2 Announce Type: cross Abstract: Adversarial scenario generation is a cost-effective approach for safety assessment of autonomous driving systems. However, existing methods are often

Zero-Shot Signal Temporal Logic Planning with Disjunctive Branch Selection in Dynamic Semantic Maps

SafetyDGX agent

arXiv:2605.01222v1 Announce Type: new Abstract: Signal Temporal Logic (STL) offers verifiable task specifications and is crucial for safety-critical control. Yet STL planning remains challenging: exac

5 May 2026

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

SafetyDGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure

SafetyDGX agent

arXiv:2605.00055v1 Announce Type: cross Abstract: We report a safety incident in a deployed multi-agent research system in which a primary AI agent installed 107 unauthorized software components, over

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

SafetyDGX agent

arXiv:2605.01451v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly being integrated into high-stakes public safety systems, including emergency call triage and dispatch decision

Causal Foundations of Collective Agency

SafetyDGX agent

arXiv:2605.00248v1 Announce Type: new Abstract: A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with c

Combining Facial Videos and Biosignals for Stress Estimation During Driving

SafetyDGX agent

arXiv:2601.04376v3 Announce Type: replace Abstract: Reliable stress recognition is critical in applications such as medical monitoring and safety-critical systems, including real-world driving. While

Edge Case Detection in Automated Driving: Methods, Challenges and Future Directions

SafetyDGX agent

arXiv:2410.08491v2 Announce Type: replace Abstract: Automated vehicles promise to enhance transportation safety and efficiency. However, ensuring their reliability in real-world conditions remains cha

Hyp2Former: Hierarchy-Aware Hyperbolic Embeddings for Open-Set Panoptic Segmentation

SafetyDGX agent

arXiv:2605.02580v1 Announce Type: new Abstract: Recognizing unknown objects is crucial for safety-critical applications such as autonomous driving and robotics. Open-Set Panoptic Segmentation (OPS) ai

Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy

SafetyDGX agent

arXiv:2509.21173v5 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have revolutionized zero-shot classification and safety-critical tasks, including Out-of-Distribution (OO

Parking Assistance for Trailer-Truck Transport Vehicles Using Sensor Fusion and Motion Planning

SafetyDGX agent

arXiv:2605.02716v1 Announce Type: new Abstract: Autonomous driving technology has rapidly evolved over the past decade, offering significant improvements in transportation efficiency, safety, and cost

UCATSC: Uncertainty-Aware Constrained Traffic Signal Control Under Vision-Based Partial Observability

SafetyDGX agent

arXiv:2602.07784v3 Announce Type: replace Abstract: Camera-based adaptive traffic signal control is inherently partially observable: detections can be missed, vehicle speeds and distances can be noisy

4 May 2026

Dynamic-TD3: A Novel Algorithm for UAV Path Planning with Dynamic Obstacle Trajectory Prediction

SafetyDGX agent

arXiv:2605.00059v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) finds extensive application in autonomous drone navigation within complex, high-risk environments. However, its practi

Impact of Task Phrasing on Presumptions in Large Language Models

SafetyDGX agent

arXiv:2605.00436v1 Announce Type: new Abstract: Concerns with the safety and reliability of applying large-language models (LLMs) in unpredictable real-world applications motivate this study, which ex

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance

SafetyDGX agent

arXiv:2605.00553v1 Announce Type: new Abstract: Large Language Model (LLM) Red-Teaming, which proactively identifies vulnerabilities of LLMs, is an essential process for ensuring safety. Finding effec

← Previous
1…2324252627…240
Next →