AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
26 May 2026

TorchLean: Formalizing Neural Networks in Lean

SafetyDGX agent

arXiv:2602.22631v2 Announce Type: replace-cross Abstract: Neural networks are increasingly deployed in scientific, safety critical, and mission critical pipelines, yet verification and analysis are of

25 May 2026

Graph-based Complexity Forecasts in UK En Route Airspace Using Relevant Aircraft Interactions

SafetyDGX agent

arXiv:2605.23696v1 Announce Type: new Abstract: Effectively managing Air Traffic Control Officer (ATCO) workload is crucial in maintaining operational safety. Group supervisors use tools that estimate

Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.23320v1 Announce Type: new Abstract: Ventilator decision support requires sequential decisions that track evolving physiology and disease trajectories while respecting safety boundaries and

People are often confused that I am against the framing of 'tool AI' This is hands down the best post explaining (some of) the issues with t…

SafetyDGX agent

People are often confused that I am against the framing of 'tool AI' This is hands down the best post explaining (some of) the issues with the term. Give it a read! Many in AI safety advocacy argue th

Relevant Walk Search for Explaining Graph Neural Networks

SafetyDGX agent

arXiv:2605.23673v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have become important machine learning tools for graph analysis, and its explainability is crucial for safety, fairness, an

The Pope rightly warns that AI must serve human dignity, not become a tool of domination or exclusion. But if we hand governments sweeping p…

SafetyDGX agent

The Pope rightly warns that AI must serve human dignity, not become a tool of domination or exclusion. But if we hand governments sweeping power over AI development in the name of safety, how do we pr

UfM*: Uncertainty from Motion* for DNN Depth Estimation Using Gaussians

SafetyDGX agent

arXiv:2605.23098v1 Announce Type: new Abstract: Reliable uncertainty estimation is critical for deploying monocular depth deep neural networks (DNNs) in safety-critical robotic systems. Conventional u

24 May 2026

even @geohotz is starting to sound like me 🤣

SafetyDGX agent

Gary Marcus humorously notes that George Hotz, an AI researcher and entrepreneur, is beginning to echo Marcus's own views or criticisms, likely regarding AI safety, limitations, or technical concerns.

23 May 2026

Expectation Consistency Loss: Rethink Confidence Calibration under Covariate Shift

SafetyDGX agent

arXiv:2605.21552v1 Announce Type: new Abstract: Confidence calibration for classification models is vital in safety-critical decision-making scenarios and has received extensive attention. General con

MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data

SafetyDGX agent

arXiv:2605.22775v1 Announce Type: new Abstract: Real-time cognitive load assessment from eye-tracking signals could potentially enable adaptive human-centered-AI such as safety-critical applications s

The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning

SafetyDGX agent

arXiv:2605.22800v1 Announce Type: new Abstract: Robustness, domain adaptation, photometric and occlusion invariance, compositional generalisation, temporal robustness, alignment safety, and classical

Visibility nowcasting in South Korea: a machine learning approach to class imbalance and distribution shift

SafetyDGX agent

arXiv:2605.21507v1 Announce Type: cross Abstract: Atmospheric visibility is a critical variable for transportation safety and air quality management, however, accurate prediction remains challenging d

22 May 2026

Bounding-Box Trajectories Matter for Video Anomaly Detection

SafetyDGX agent

arXiv:2605.21957v1 Announce Type: new Abstract: Video anomaly detection is critical for public safety and security, yet remains highly challenging despite extensive research due to large variations in

LACO: Adaptive Latent Communication for Collaborative Driving

SafetyDGX agent

arXiv:2605.22504v1 Announce Type: cross Abstract: Collaborative driving aims to improve safety and efficiency by enabling connected vehicles to coordinate under partial observability. Recent approache

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

SafetyDGX agent

arXiv:2605.21168v1 Announce Type: new Abstract: Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress test

21 May 2026

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards

SafetyDGX agent

arXiv:2605.21180v1 Announce Type: new Abstract: Large language models show strong potential for automated code generation, but lack guarantees for correctness, quality, safety, and domain-specific con

LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models

SafetyDGX agent

arXiv:2605.21362v1 Announce Type: new Abstract: Jailbreak attacks expose a persistent gap between the intended safety behavior of aligned large language models and their behavior under adversarial pro

Proximal State Nudging: Reducing Skill Atrophy from AI Assistance

SafetyDGX agent

arXiv:2605.20355v1 Announce Type: cross Abstract: Skill atrophy, the gradual decline of human capability under AI assistance, poses a safety risk in shared-control of semi-autonomous systems, where op

Shipping features to production just got easier with new feature flags in AppLifecycle Manager

SafetyDGX agent

Many development teams are familiar with the hesitation that comes right before pushing a new feature live. As AI helps developers write code faster, the gap between rapid code generation and safe pro

20 May 2026

Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models

SafetyDGX agent

arXiv:2410.15362v2 Announce Type: replace-cross Abstract: Aligned Large Language Models (LLMs) have attracted significant attention for their safety, particularly in the context of jailbreak attacks t

FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models

SafetyDGX agent

arXiv:2605.19739v1 Announce Type: new Abstract: Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the

Hard-Label Black-Box Attacks on 3D Point Clouds

SafetyDGX agent

arXiv:2412.00404v2 Announce Type: replace Abstract: With the maturity of depth sensors in various 3D safety-critical applications, 3D point cloud models have been shown to be vulnerable to adversarial

Implicit Action Chunking for Smooth Continuous Control

SafetyDGX agent

arXiv:2605.19592v1 Announce Type: cross Abstract: Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deploym

Improved visual-information-driven model for crowd simulation and its modular application

SafetyDGX agent

arXiv:2504.03758v4 Announce Type: replace-cross Abstract: Crowd movement simulation is crucial for pedestrian safety management and facility design. Data-driven models offer the potential to improve r

Sampling-Based Safe Reinforcement Learning

SafetyDGX agent

arXiv:2605.19469v1 Announce Type: cross Abstract: Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sa

19 May 2026

Activation Steering with a Feedback Controller

SafetyDGX agent

arXiv:2510.04309v3 Announce Type: replace Abstract: Controlling the behaviors of large language models (LLM) is fundamental to their safety alignment and reliable deployment. However, existing steerin

Assessing Localization Technologies for Pedestrian Collision Avoidance

SafetyDGX agent

arXiv:2605.18295v1 Announce Type: new Abstract: Robust pedestrian safety is crucial to the next-generation of intelligent transportation systems. Such systems rely on active pedestrian localization an

Assured autonomy: How operations research powers and orchestrates generative AI systems

SafetyDGX agent

arXiv:2512.23978v2 Announce Type: replace Abstract: Generative artificial intelligence (GenAI) is shifting from conversational assistants toward agentic systems -- autonomous decision-making systems t

Constrained Policy Optimization via Sampling-Based Weight-Space Projection

Model ReleasesDGX agent

arXiv:2512.13788v2 Announce Type: replace Abstract: Safety-critical learning requires policies that improve performance without leaving the safe operating regime. We study constrained policy learning

Density-Ratio Weighted Behavioral Cloning: Learning Control Policies from Corrupted Datasets

SafetyDGX agent

arXiv:2510.01479v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables policy optimization from fixed datasets, making it suitable for safety-critical applications where onlin

Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking

SafetyDGX agent

arXiv:2506.12617v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in human-AI interaction research and practice, yet existing capability and safety benchmarks reve

Forgetting is Competition: Rethinking Unlearning as Representation Interference in Diffusion Models

SafetyDGX agent

arXiv:2603.00975v2 Announce Type: replace-cross Abstract: Deployed text-to-image diffusion models increasingly require post-hoc concept unlearning for copyright claims, artist opt-outs, safety updates

ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization

SafetyDGX agent

arXiv:2605.18320v1 Announce Type: cross Abstract: Offline reinforcement learning methods typically enforce strict constraints to ensure safety; yet this rigidity often prevents the discovery of optima

Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems

SafetyDGX agent

arXiv:2605.16278v1 Announce Type: cross Abstract: The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that ma

Meltdown: Circuits and Bifurcations in Point-Cloud-Conditioned 3D Diffusion Transformers

SafetyDGX agent

arXiv:2602.11130v2 Announce Type: replace-cross Abstract: Sparse point clouds are a common input modality for 3D surface reconstruction, including in safety-critical settings such as surgical navigati

Semantic Smoothing via Novel View Synthesis for Robust SAR Image Classification

SafetyDGX agent

arXiv:2605.16440v1 Announce Type: cross Abstract: Deep neural networks are vulnerable to adversarial perturbations, limiting deployment in safety-critical applications such as synthetic aperture radar

Uncertainty Reliability Under Domain Shift: An Investigation for Data-Driven Blood Pressure Estimation in Photoplethysmography

SafetyDGX agent

arXiv:2605.18008v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is critical for safety-critical domains like healthcare, yet it is rarely evaluated under realistic out-of-distribution

18 May 2026

3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation

Model ReleasesDGX agent

arXiv:2605.15398v1 Announce Type: cross Abstract: Recent advances in 3D generative editing, particularly pipelines based on 3D Gaussian Splatting (3DGS), have achieved high-fidelity, multi-view-consis

Learning Context-conditioned Gaussian Overbounds for Convolution-Based Uncertainty Propagation

SafetyDGX agent

arXiv:2605.15789v1 Announce Type: new Abstract: Uncertainty quantification is essential in safety-critical settings--from autonomous driving to aviation, finance, and health--where decisions must rely

Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise

SafetyDGX agent

arXiv:2605.15641v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general planners in embodied intelligence, enabling high level coordination and low level task pla

Towards a more realistic evaluation of machine learning models for bearing fault diagnosis

SafetyDGX agent

arXiv:2509.22267v4 Announce Type: replace Abstract: Reliable detection of bearing faults is essential for maintaining the safety and operational efficiency of rotating machinery. While recent advances

17 May 2026

This was always and only the sole and exclusive purpose of the UK online censorship act, to help Labour nuke its enemies

SafetyDGX agent

This was always and only the sole and exclusive purpose of the UK online censorship act, to help Labour nuke its enemies 🚨 Labour is using the “Online Safety Act” to silence political opponents, and T

15 May 2026

Do Reasoning LLMs Refuse What They Infer in Long Contexts?

SafetyDGX agent

arXiv:2602.08874v2 Announce Type: replace Abstract: Long-context LLMs can infer objectives that are not stated explicitly. This capability is useful for reasoning over documents, code, retrieved evide

Exploring Geographic Relative Space in Large Language Models through Activation Patching

SafetyDGX agent

arXiv:2605.14535v1 Announce Type: new Abstract: The increased use of Large Language Models (LLMs) in geography raises substantial questions about the safety of integrating these tools across a wide ra

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries

SafetyDGX agent

arXiv:2605.14605v1 Announce Type: cross Abstract: Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs. Although these models are safety-aligned

Precise Verification of Transformers through ReLU-Catalyzed Abstraction Refinement

SafetyDGX agent

arXiv:2605.14294v1 Announce Type: new Abstract: Formal verification of transformers has become increasingly important due to their widespread deployment in safety-critical applications. Compared to cl

Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion

SafetyDGX agent

arXiv:2605.14396v1 Announce Type: new Abstract: Autonomous vehicles depend on online HD map construction to perceive lane boundaries, dividers, and pedestrian crossings -- safety-critical road element

14 May 2026

A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline

SafetyDGX agent

arXiv:2605.12608v1 Announce Type: new Abstract: Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains

Belief-Space Residual Risk for Automated Driving under Localization Uncertainty

SafetyDGX agent

arXiv:2605.12710v1 Announce Type: new Abstract: Residual risk metrics have recently been introduced to assess the safety implications of automated driving systems. Existing approaches typically assume

Digital Twins as Synthetic Controls in Single-Arm Trials

SafetyDGX agent

arXiv:2605.12832v1 Announce Type: cross Abstract: Single-arm trials are an important study design for evaluating drug efficacy and safety without enrolling patients into a control arm. Although they d

Humanwashing -- It Should Leave You Feeling Dirty

SafetyDGX agent

arXiv:2605.13723v1 Announce Type: cross Abstract: The phrase 'human in the loop' is increasingly used to imply a sense of safety in relation to AI decision systems. It shouldn't. There are contexts wh

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

SafetyDGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

NeuroRisk: Physics-Informed Neural Optimization for Risk-Aware Traffic Engineering

SafetyDGX agent

arXiv:2605.12862v1 Announce Type: cross Abstract: In production Wide-Area Networks (WANs), correlated failures dominate availability losses, forcing operators to reserve large safety margins that leav

Quantifying Sensitivity for Tree Ensembles: A symbolic and compositional approach

SafetyDGX agent

arXiv:2605.13830v1 Announce Type: new Abstract: Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence verif

Quantitative Certification of Agentic Tool Selection

SafetyDGX agent

arXiv:2510.03992v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant extern

SPOT: Selective Prompt Projection via Total Variation for Inference-Only Safe Text-to-Image Generation

SafetyDGX agent

arXiv:2602.00616v3 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models enable high quality open ended synthesis, but practical use requires suppressing unsafe generations while prese

Watermarking Should Be Treated as a Monitoring Primitive

SafetyDGX agent

arXiv:2605.13095v1 Announce Type: cross Abstract: Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under adversa

13 May 2026

Few-Shot Synthetic Data Generation with Diffusion Models for Downstream Vision Tasks

SafetyDGX agent

arXiv:2605.11898v1 Announce Type: new Abstract: Class imbalance is a persistent challenge in visual recognition, particularly in safety-critical domains where collecting positive examples is expensive

Interpreting Context-Aware Human Preferences for Multi-Objective Robot Navigation

SafetyDGX agent

arXiv:2603.17510v2 Announce Type: replace Abstract: Robots operating in human-shared environments must not only achieve task-level navigation objectives such as safety and efficiency, but also adapt t

Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation

SafetyDGX agent

arXiv:2605.11730v1 Announce Type: new Abstract: Automated red-teaming for LLMs often discovers narrow attack slices, missing diverse real-world threats, and yielding insufficient data for safety fine-

← Previous
1…2223242526…240
Next →