AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management

DGX agent

arXiv:2604.21209v1 Announce Type: new Abstract: Online reviews have played a pivotal role in consumers' decision-making processes. Existing research has highlighted the significant impact of manageria

safetyarxiv-cs-ai
24 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

And eventually schools will (if they have any sense) probably withdraw or limit LLMs.

DGX agent

And eventually schools will (if they have any sense) probably withdraw or limit LLMs. 🦔Schools across the US are reversing years of technology-first classroom policies after studies show laptop and sc

safetygary-marcus--x
24 Apr 2026
Safety

Clinical Reasoning AI for Oncology Treatment Planning: A Multi-Specialty Case-Based Evaluation

DGX agent

arXiv:2604.20869v1 Announce Type: cross Abstract: Background: More than 80% of U.S. cancer care is delivered in community settings, where survival remains worse than at academic centers. Clinicians mu

safetyarxiv-cs-ai
24 Apr 2026
Safety

FryNet: Dual-Stream Adversarial Fusion for Non-Destructive Frying Oil Oxidation Assessment

DGX agent

arXiv:2604.21321v1 Announce Type: new Abstract: Monitoring frying oil degradation is critical for food safety, yet current practice relies on destructive wet-chemistry assays that provide no spatial i

safetyarxiv-cs-cv
24 Apr 2026
Safety

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs

DGX agent

arXiv:2604.21926v1 Announce Type: new Abstract: Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in priva

safetyarxiv-cs-cv
24 Apr 2026
Safety

Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models

DGX agent

arXiv:2604.20995v1 Announce Type: new Abstract: Alignment faking, where a model behaves aligned with developer policy when monitored but reverts to its own preferences when unobserved, is a concerning

safetyarxiv-cs-ai
24 Apr 2026
Safety

DAIRE: A lightweight AI model for real-time detection of Controller Area Network attacks in the Internet of Vehicles

DGX agent

arXiv:2604.20771v1 Announce Type: cross Abstract: The Internet of Vehicles (IoV) is advancing modern transportation by improving safety, efficiency, and intelligence. However, the reliance on the Cont

safetyarxiv-cs-ai
23 Apr 2026
Safety

Diagnosing CFG Interpretation in LLMs

DGX agent

arXiv:2604.20811v1 Announce Type: new Abstract: As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as

safetyarxiv-cs-ai
23 Apr 2026
Safety

Evaluating Assurance Cases as Text-Attributed Graphs for Structure and Provenance Analysis

DGX agent

arXiv:2604.20577v1 Announce Type: cross Abstract: An assurance case is a structured argument document that justifies claims about a system's requirements or properties, which are supported by evidence

safetyarxiv-cs-lg
23 Apr 2026
Safety

Generative Augmentation of Imbalanced Flight Records for Flight Diversion Prediction: A Multi-objective Optimisation Framework

DGX agent

arXiv:2604.20288v1 Announce Type: new Abstract: Flight diversions are rare but high-impact events in aviation, making their reliable prediction vital for both safety and operational efficiency. Howeve

safetyarxiv-cs-lg
23 Apr 2026
Safety

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

DGX agent

arXiv:2604.20382v1 Announce Type: new Abstract: Rising demand for mental health support has increased interest in using Large Language Models (LLMs) for counseling. However, adapting LLMs to this high

safetyarxiv-cs-cl
23 Apr 2026
Safety

More people will die from suppressing AI than from the imaginary AI apocalypse. They'll die from restricting safe self-driving cars that are…

DGX agent

More people will die from suppressing AI than from the imaginary AI apocalypse. They'll die from restricting safe self-driving cars that are 90% better drivers than people who kill 1.5 million people

safetyyann-lecun--x
23 Apr 2026
Safety

Toward Cooperative Driving in Mixed Traffic: An Adaptive Potential Game-Based Approach with Field Test Verification

DGX agent

arXiv:2604.20231v1 Announce Type: new Abstract: Connected autonomous vehicles (CAVs), which represent a significant advancement in autonomous driving technology, have the potential to greatly increase

safetyarxiv-cs-ro
23 Apr 2026
Safety

AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos

DGX agent

arXiv:2604.18993v1 Announce Type: cross Abstract: Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-w

safetyarxiv-cs-ai
22 Apr 2026
Safety

Benchmarking Misuse Mitigation Against Covert Adversaries

DGX agent

arXiv:2506.06414v2 Announce Type: replace-cross Abstract: Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing sa

safetyarxiv-cs-ai
22 Apr 2026
Safety

Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

DGX agent

arXiv:2510.18358v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensemb

safetyarxiv-cs-cv
22 Apr 2026
Safety

From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing

DGX agent

arXiv:2604.18918v1 Announce Type: cross Abstract: Simulation-based testing of autonomous driving systems (ADS) must uncover realistic and diverse failures in dense, heterogeneous traffic. However, exi

safetyarxiv-cs-lg
22 Apr 2026
Safety

Mind2Drive: Predicting Driver Intentions from EEG in Real-world On-Road Driving

DGX agent

arXiv:2604.19368v1 Announce Type: new Abstract: Predicting driver intention from neurophysiological signals offers a promising pathway for enhancing proactive safety in advanced driver assistance syst

safetyarxiv-cs-cv
22 Apr 2026
Safety

REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

DGX agent

arXiv:2604.18757v1 Announce Type: cross Abstract: The retina provides a unique, noninvasive window into Alzheimer's disease (AD) and dementia, capturing early structural changes through morphometric f

safetyarxiv-cs-ai
22 Apr 2026
Safety

When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide

DGX agent

arXiv:2604.19206v1 Announce Type: new Abstract: The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely ha

safetyarxiv-cs-cv
22 Apr 2026
Safety

Alignment Data Map for Efficient Preference Data Selection and Diagnosis

DGX agent

arXiv:2505.23114v3 Announce Type: replace Abstract: Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and ineffic

safetyarxiv-cs-cl
21 Apr 2026
Safety

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

DGX agent

arXiv:2604.18468v1 Announce Type: new Abstract: Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before rea

safetyarxiv-cs-cv
21 Apr 2026
Safety

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs

DGX agent

arXiv:2511.02356v2 Announce Type: replace-cross Abstract: Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. However, existing methods generally l

safetyarxiv-cs-lg
21 Apr 2026
Safety

Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning

DGX agent

arXiv:2604.16702v1 Announce Type: new Abstract: Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestria

safetyarxiv-cs-ro
21 Apr 2026
Safety

Characterizing Model-Native Skills

DGX agent

arXiv:2604.17614v1 Announce Type: cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on

safetyarxiv-cs-cl
21 Apr 2026
Safety

Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation

DGX agent

arXiv:2604.17178v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) plays a critical role in mental health assistance by providing accessible psychological support in real-world appli

safetyarxiv-cs-cl
21 Apr 2026
Safety

Detecting Alarming Student Verbal Responses using Text and Audio Classifier

DGX agent

arXiv:2604.16717v1 Announce Type: new Abstract: This paper addresses a critical safety gap in the use Automated Verbal Response Scoring (AVRS). We present a novel hybrid framework for troubled student

safetyarxiv-cs-cl
21 Apr 2026
Local Ai

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks

DGX agent

arXiv:2604.18510v1 Announce Type: cross Abstract: Open-weight language models can be rendered unsafe through several distinct interventions, but the resulting models may differ substantially in capabi

local-aiarxiv-cs-cl
21 Apr 2026
Safety

Efficient Diffusion Models under Nonconvex Equality and Inequality constraints via Landing

DGX agent

arXiv:2604.17838v1 Announce Type: new Abstract: Generative modeling within constrained sets is essential for scientific and engineering applications involving physical, geometric, or safety requiremen

safetyarxiv-cs-lg
21 Apr 2026
Safety

Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models

DGX agent

arXiv:2604.16481v1 Announce Type: new Abstract: Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable

safetyarxiv-cs-cv
21 Apr 2026
Safety

HAVEN: Hierarchical Adversary-aware Visibility-Enabled Navigation with Cover Utilization using Deep Transformer Q-Networks

DGX agent

arXiv:2512.00592v2 Announce Type: replace Abstract: Autonomous navigation in partially observable environments requires agents to reason beyond immediate sensor input, exploit occlusion, and ensure sa

safetyarxiv-cs-ro
21 Apr 2026
Safety

Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization

DGX agent

arXiv:2604.17984v1 Announce Type: new Abstract: Uncertainty quantification is crucial in safety-critical systems, where decisions must be made under uncertainty. In particular, we consider the problem

safetyarxiv-cs-lg
21 Apr 2026
Safety

The Impact of Off-Policy Training Data on Probe Generalisation

DGX agent

arXiv:2511.17408v4 Announce Type: replace-cross Abstract: Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning b

safetyarxiv-cs-lg
21 Apr 2026
Safety

They just indicted the Southern Poverty Law Center on 11 counts of fraud. Wow! The SPLC worked with Facebook, Old Twitter and Youtube on the…

DGX agent

They just indicted the Southern Poverty Law Center on 11 counts of fraud. Wow! The SPLC worked with Facebook, Old Twitter and Youtube on their safety teams, sometimes directly advising/flagging what c

safetyelon-musk--x
21 Apr 2026
Safety

AI Agents and Hard Choices

DGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

safetyarxiv-cs-ai
20 Apr 2026
Safety

Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover

DGX agent

arXiv:2603.11331v2 Announce Type: replace-cross Abstract: Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. Empirically, we find that strong adversari

safetyarxiv-cs-ai
20 Apr 2026
Safety

Safe and Energy-Aware Multi-Robot Density Control via PDE-Constrained Optimization for Long-Duration Autonomy

DGX agent

arXiv:2604.15524v1 Announce Type: cross Abstract: This paper presents a novel density control framework for multi-robot systems with spatial safety and energy sustainability guarantees. Stochastic rob

safetyarxiv-cs-ro
20 Apr 2026
Safety

Safe Deep Reinforcement Learning for Building Heating Control and Demand-side Flexibility

DGX agent

arXiv:2604.16033v1 Announce Type: cross Abstract: Buildings account for approximately 40% of global energy consumption, and with the growing share of intermittent renewable energy sources, enabling de

safetyarxiv-cs-ai
20 Apr 2026
Safety

Dario is not as different from Sam as people seem to think.

DGX agent

Gary Marcus argues that Dario Amodei (Anthropic CEO) and Sam Altman (OpenAI CEO) share more similarities in their approaches to AI development and safety than commonly perceived, despite their public

safetygary-marcus--x
18 Apr 2026
Safety

Context Over Content: Exposing Evaluation Faking in Automated Judges

DGX agent

arXiv:2604.15224v1 Announce Type: cross Abstract: The extit{LLM-as-a-judge} paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: th

safetyarxiv-cs-cl
17 Apr 2026
Safety

[Emerging Ideas] Artificial Tripartite Intelligence: A Bio-Inspired, Sensor-First Architecture for Physical AI

DGX agent

arXiv:2604.13959v1 Announce Type: new Abstract: As AI moves from data centers to robots and wearables, scaling ever-larger models becomes insufficient. Physical AI operates under tight latency, energy

safetyarxiv-cs-ai
17 Apr 2026
Safety

From Risk to Rescue: An Agentic Survival Analysis Framework for Liquidation Prevention

DGX agent

arXiv:2604.14583v1 Announce Type: new Abstract: Decentralized Finance (DeFi) lending protocols like Aave v3 rely on over-collateralization to secure loans, yet users frequently face liquidation due to

safetyarxiv-cs-lg
17 Apr 2026
Safety

Integrating Object Detection, LiDAR-Enhanced Depth Estimation, and Segmentation Models for Railway Environments

DGX agent

arXiv:2604.14781v1 Announce Type: new Abstract: Obstacle detection in railway environments is crucial for ensuring safety. However, very few studies address the problem using a complete, modular, and

safetyarxiv-cs-cv
17 Apr 2026
Safety

Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees

DGX agent

arXiv:2604.14243v1 Announce Type: new Abstract: Real-world decision-making systems operate in environments where state transitions depend not only on the agent's actions, but also on extbf{exogenous f

safetyarxiv-cs-lg
17 Apr 2026
Safety

The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious

DGX agent

arXiv:2604.14414v1 Announce Type: new Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However

safetyarxiv-cs-cl
17 Apr 2026
Safety

C^2T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination

DGX agent

arXiv:2604.13098v1 Announce Type: cross Abstract: State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (

safetyarxiv-cs-cv
16 Apr 2026
Safety

From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution

DGX agent

arXiv:2604.14053v1 Announce Type: new Abstract: Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves infe

safetyarxiv-cs-cl
16 Apr 2026
Safety

Learning Probabilistic Responsibility Allocations for Multi-Agent Interactions

DGX agent

arXiv:2604.13128v1 Announce Type: cross Abstract: Human behavior in interactive settings is shaped not only by individual objectives but also by shared constraints with others, such as safety. Underst

safetyarxiv-cs-lg
16 Apr 2026
← Previous
1…3132333435…299
Next →