AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
4 May 2026

Unlearning What Matters: Token-Level Attribution for Precise Language Model Unlearning

SafetyDGX agent

arXiv:2605.00364v1 Announce Type: new Abstract: Machine unlearning has emerged as a critical capability for addressing privacy, safety, and regulatory concerns in large language models (LLMs). Existin

1 May 2026

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

SafetyDGX agent

arXiv:2604.27019v1 Announce Type: cross Abstract: Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this t

Efficient Preimage Approximation for Neural Network Certification

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2505.22798v3 Announce Type: replace-cross Abstract: The growing reliance on artificial intelligence in safety- and security-critical applications is raising concerns about the robustness of neur

Elon Musk, Sam Altman, the future of humanity, and … goblins. Terrific interview @theinformation with @rocketalignment https://www.youtube.c…

SafetyDGX agent

This post references an interview featuring Elon Musk and Sam Altman discussing AI safety, existential risks, and the future of humanity, with an unconventional or humorous element involving goblins.

OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment

Model ReleasesDGX agent

arXiv:2506.22500v2 Announce Type: replace-cross Abstract: Automated identification of surgical safety risks is critical for improving patient outcomes; however, Multimodal Large Language Models (MLLMs

The Field of Safe Motion: Operationalizing Affordances in the Field of Safe Travel Using Reachability Analysis

SafetyDGX agent

arXiv:2604.27168v1 Announce Type: new Abstract: We present the Field of Safe Motion (FSM), a quantitative safety model for determining whether a driver maintains a collision-free escape route, or 'out

30 Apr 2026

A Scaled Three-Vehicle Platooning Platform

SafetyDGX agent

arXiv:2604.25963v1 Announce Type: new Abstract: Vehicle platooning has attracted increasing attention as a promising approach to improve traffic efficiency, energy consumption, and roadway safety thro

Adversarial Robustness of NTK Neural Networks

SafetyDGX agent

arXiv:2604.25965v1 Announce Type: cross Abstract: Deep learning models are widely deployed in safety-critical domains, but remain vulnerable to adversarial attacks. In this paper, we study the adversa

Crime Hotspot Prediction Using Deep Graph Convolutional Networks

SafetyDGX agent

arXiv:2506.13116v2 Announce Type: replace-cross Abstract: Crime hotspot prediction is critical for ensuring urban safety and effective law enforcement, it remains challenging due to complex spatial de

Improving Bayesian Optimization for Portfolio Management with an Adaptive Scheduling

SafetyDGX agent

arXiv:2504.13529v4 Announce Type: replace Abstract: Existing black-box portfolio management systems are prevalent in the financial industry due to commercial and safety constraints, though their perfo

Uncertainty-Aware Information Pursuit for Interpretable and Reliable Medical Image Analysis

SafetyDGX agent

arXiv:2506.16742v3 Announce Type: replace Abstract: To be adopted in safety-critical domains like medical image analysis, AI systems must provide human-interpretable decisions. Variational Information

29 Apr 2026

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

SafetyDGX agent

arXiv:2604.25203v1 Announce Type: new Abstract: Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs

Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models

SafetyDGX agent

arXiv:2512.20677v4 Announce Type: replace-cross Abstract: The increasing deployment of large language models (LLMs) in safety-critical applications raises fundamental challenges in systematically eval

RISE: Self-Improving Robot Policy with Compositional World Model

SafetyDGX agent

arXiv:2602.11075v2 Announce Type: replace Abstract: Despite the sustained scaling on model capacity and data acquisition, Vision-Language-Action (VLA) models remain brittle in contact-rich and dynamic

Scary that he doesn’t know this.

SafetyDGX agent

Scary that he doesn’t know this. Elon Musk just said on the witness stand that he doesn't know what a 'safety card' is with regard to AI development... 'Why would it be a card?' He also says he doesn'

'The more you look at chatbots, the more you realize that they were rushed to market with very little consideration for the consequences,' h…

SafetyDGX agent

Gary Marcus critiques the rapid deployment of chatbots to market without adequate consideration of potential harms and consequences. The statement reflects concerns about insufficient safety testing a

28 Apr 2026

An Automatic Ground Collision Avoidance System with Reinforcement Learning

SafetyDGX agent

arXiv:2604.24403v1 Announce Type: new Abstract: This article evaluates an artificial intelligence (AI)-based Automatic Ground Collision Avoidance System (AGCAS) designed for advanced jet trainers to e

An Information-Geometric Framework for Stability Analysis of Large Language Models under Entropic Stress

SafetyDGX agent

arXiv:2604.24076v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in high-stakes and operational settings, evaluation strategies based solely on aggregate accur

Beyond Binary Out-of-Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories

SafetyDGX agent

arXiv:2510.17381v2 Announce Type: replace Abstract: Detecting out-of-distribution (OOD) data is critical for machine learning, be it for safety reasons or to enable open-ended learning. However, beyon

Cooperative Informative Sensing for Monitoring Dynamic Indoor Environments via Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2604.23179v1 Announce Type: cross Abstract: Monitoring human activity in indoor environments is important for applications such as facility management, safety assessment, and space utilization a

Extending Precipitation Nowcasting Horizons via Spectral Fusion of Radar Observations and Foundation Model Priors

SafetyDGX agent

arXiv:2603.21768v3 Announce Type: replace-cross Abstract: Precipitation nowcasting is critical for disaster mitigation and aviation safety. However, radar-only models frequently suffer from a lack of

honestly, one of the worst things this nation has ever done to itself.

SafetyDGX agent

honestly, one of the worst things this nation has ever done to itself. No amount of words can convey how far behind this country will be in every avenue for a number of years because of this administr

I spent the last 2 months trying to prevent this. If OpenAI offered a fig leaf, Google said 'imagine we offered a fig leaf.' Google affirms …

SafetyDGX agent

I spent the last 2 months trying to prevent this. If OpenAI offered a fig leaf, Google said 'imagine we offered a fig leaf.' Google affirms it can't veto usage, commits to modify safety filters at gov

IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance

SafetyDGX agent

arXiv:2604.23446v1 Announce Type: new Abstract: Industrial maintenance environments increasingly rely on AI systems to assist operators in understanding asset behavior, diagnosing failures, and evalua

iWatchRoad: Scalable Detection and Geospatial Visualization of Potholes for Smart Cities

SafetyDGX agent

arXiv:2508.10945v2 Announce Type: replace Abstract: Potholes on the roads are a serious hazard and maintenance burden. This poses a significant threat to road safety and vehicle longevity, especially

Large Language Model based Interactive Decision-Making for Autonomous Driving

SafetyDGX agent

arXiv:2604.23513v1 Announce Type: new Abstract: In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly cons

27 Apr 2026

On Star Price's new @hitpausepod podcast, ControlAI's US Director Connor Leahy (@NPCollapse) explains how AIs can already tell when they're …

SafetyDGX agent

On Star Price's new @hitpausepod podcast, ControlAI's US Director Connor Leahy (@NPCollapse) explains how AIs can already tell when they're being tested. Currently, we can still catch them cheating, b

Recognition Without Authorization: LLMs and the Moral Order of Online Advice

SafetyDGX agent

arXiv:2604.22143v1 Announce Type: cross Abstract: Large language models are increasingly used to mediate everyday interpersonal dilemmas, yet how their advisory defaults interact with the concentrated

RedVLA: Physical Red Teaming for Vision-Language-Action Models

SafetyDGX agent

arXiv:2604.22591v1 Announce Type: new Abstract: The real-world deployment of Vision-Language-Action (VLA) models remains limited by the risk of unpredictable and irreversible physical harm. However, w

The EU AI Act is the first comprehensive regulation for AI systems, and it’s right around the corner. With the clock running down, we took a…

SafetyDGX agent

The EU AI Act is the first comprehensive regulation for AI systems, and it’s right around the corner. With the clock running down, we took a deeper look at what these requirements mean for your agents

This is the 7th consecutive week with new reporting that ChatGPT was used in connection with murder or suicide. OpenAI doesn’t “benefit all …

SafetyDGX agent

This is the 7th consecutive week with new reporting that ChatGPT was used in connection with murder or suicide. OpenAI doesn’t “benefit all of humanity.” Safety researchers keep quitting OpenAI becaus

Transferable Physical-World Adversarial Patches Against Pedestrian Detection Models

SafetyDGX agent

arXiv:2604.22552v1 Announce Type: new Abstract: Physical adversarial patch attacks critically threaten pedestrian detection, causing surveillance and autonomous driving systems to miss pedestrians and

V-STC: A Time-Efficient Multi-Vehicle Coordinated Trajectory Planning Approach

SafetyDGX agent

arXiv:2604.22196v1 Announce Type: new Abstract: Coordinating the motions of multiple autonomous vehicles (AVs) requires planning frameworks that ensure safety while making efficient use of space and t

26 Apr 2026

The only people who believe any of this are non-coders. I tried to build a game (an area I’m an n00b in.) The results are amusingly disastro…

SafetyDGX agent

The only people who believe any of this are non-coders. I tried to build a game (an area I’m an n00b in.) The results are amusingly disastrous - I never before coded a decent game. But I’ll crack out

25 Apr 2026

David Friedberg on the Nonprofit Scam: 90% Are Bullsh*t “ The definition of exempt activities is charitable, religious, educational, scienti…

SafetyDGX agent

David Friedberg on the Nonprofit Scam: 90% Are Bullsh*t “ The definition of exempt activities is charitable, religious, educational, scientific, literacy, public safety, or fostering amateur sports co

24 Apr 2026

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management

SafetyDGX agent

arXiv:2604.21209v1 Announce Type: new Abstract: Online reviews have played a pivotal role in consumers' decision-making processes. Existing research has highlighted the significant impact of manageria

And eventually schools will (if they have any sense) probably withdraw or limit LLMs.

SafetyDGX agent

And eventually schools will (if they have any sense) probably withdraw or limit LLMs. 🦔Schools across the US are reversing years of technology-first classroom policies after studies show laptop and sc

Clinical Reasoning AI for Oncology Treatment Planning: A Multi-Specialty Case-Based Evaluation

SafetyDGX agent

arXiv:2604.20869v1 Announce Type: cross Abstract: Background: More than 80% of U.S. cancer care is delivered in community settings, where survival remains worse than at academic centers. Clinicians mu

FryNet: Dual-Stream Adversarial Fusion for Non-Destructive Frying Oil Oxidation Assessment

SafetyDGX agent

arXiv:2604.21321v1 Announce Type: new Abstract: Monitoring frying oil degradation is critical for food safety, yet current practice relies on destructive wet-chemistry assays that provide no spatial i

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs

SafetyDGX agent

arXiv:2604.21926v1 Announce Type: new Abstract: Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in priva

Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models

SafetyDGX agent

arXiv:2604.20995v1 Announce Type: new Abstract: Alignment faking, where a model behaves aligned with developer policy when monitored but reverts to its own preferences when unobserved, is a concerning

23 Apr 2026

DAIRE: A lightweight AI model for real-time detection of Controller Area Network attacks in the Internet of Vehicles

SafetyDGX agent

arXiv:2604.20771v1 Announce Type: cross Abstract: The Internet of Vehicles (IoV) is advancing modern transportation by improving safety, efficiency, and intelligence. However, the reliance on the Cont

Diagnosing CFG Interpretation in LLMs

SafetyDGX agent

arXiv:2604.20811v1 Announce Type: new Abstract: As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as

Evaluating Assurance Cases as Text-Attributed Graphs for Structure and Provenance Analysis

SafetyDGX agent

arXiv:2604.20577v1 Announce Type: cross Abstract: An assurance case is a structured argument document that justifies claims about a system's requirements or properties, which are supported by evidence

Generative Augmentation of Imbalanced Flight Records for Flight Diversion Prediction: A Multi-objective Optimisation Framework

SafetyDGX agent

arXiv:2604.20288v1 Announce Type: new Abstract: Flight diversions are rare but high-impact events in aviation, making their reliable prediction vital for both safety and operational efficiency. Howeve

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

SafetyDGX agent

arXiv:2604.20382v1 Announce Type: new Abstract: Rising demand for mental health support has increased interest in using Large Language Models (LLMs) for counseling. However, adapting LLMs to this high

More people will die from suppressing AI than from the imaginary AI apocalypse. They'll die from restricting safe self-driving cars that are…

SafetyDGX agent

More people will die from suppressing AI than from the imaginary AI apocalypse. They'll die from restricting safe self-driving cars that are 90% better drivers than people who kill 1.5 million people

Toward Cooperative Driving in Mixed Traffic: An Adaptive Potential Game-Based Approach with Field Test Verification

SafetyDGX agent

arXiv:2604.20231v1 Announce Type: new Abstract: Connected autonomous vehicles (CAVs), which represent a significant advancement in autonomous driving technology, have the potential to greatly increase

22 Apr 2026

AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos

SafetyDGX agent

arXiv:2604.18993v1 Announce Type: cross Abstract: Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-w

Benchmarking Misuse Mitigation Against Covert Adversaries

SafetyDGX agent

arXiv:2506.06414v2 Announce Type: replace-cross Abstract: Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing sa

Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

SafetyDGX agent

arXiv:2510.18358v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensemb

From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing

SafetyDGX agent

arXiv:2604.18918v1 Announce Type: cross Abstract: Simulation-based testing of autonomous driving systems (ADS) must uncover realistic and diverse failures in dense, heterogeneous traffic. However, exi

Mind2Drive: Predicting Driver Intentions from EEG in Real-world On-Road Driving

SafetyDGX agent

arXiv:2604.19368v1 Announce Type: new Abstract: Predicting driver intention from neurophysiological signals offers a promising pathway for enhancing proactive safety in advanced driver assistance syst

REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

SafetyDGX agent

arXiv:2604.18757v1 Announce Type: cross Abstract: The retina provides a unique, noninvasive window into Alzheimer's disease (AD) and dementia, capturing early structural changes through morphometric f

When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide

SafetyDGX agent

arXiv:2604.19206v1 Announce Type: new Abstract: The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely ha

21 Apr 2026

Alignment Data Map for Efficient Preference Data Selection and Diagnosis

SafetyDGX agent

arXiv:2505.23114v3 Announce Type: replace Abstract: Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and ineffic

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

SafetyDGX agent

arXiv:2604.18468v1 Announce Type: new Abstract: Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before rea

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs

SafetyDGX agent

arXiv:2511.02356v2 Announce Type: replace-cross Abstract: Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. However, existing methods generally l

Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning

SafetyDGX agent

arXiv:2604.16702v1 Announce Type: new Abstract: Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestria

Characterizing Model-Native Skills

SafetyDGX agent

arXiv:2604.17614v1 Announce Type: cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on

← Previous
1…2425262728…240
Next →