AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

DGX agent

arXiv:2604.18757v1 Announce Type: cross Abstract: The retina provides a unique, noninvasive window into Alzheimer's disease (AD) and dementia, capturing early structural changes through morphometric f

safetyarxiv-cs-ai
22 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide

DGX agent

arXiv:2604.19206v1 Announce Type: new Abstract: The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely ha

safetyarxiv-cs-cv
22 Apr 2026
Safety

Alignment Data Map for Efficient Preference Data Selection and Diagnosis

DGX agent

arXiv:2505.23114v3 Announce Type: replace Abstract: Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and ineffic

safetyarxiv-cs-cl
21 Apr 2026
Safety

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

DGX agent

arXiv:2604.18468v1 Announce Type: new Abstract: Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before rea

safetyarxiv-cs-cv
21 Apr 2026
Safety

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs

DGX agent

arXiv:2511.02356v2 Announce Type: replace-cross Abstract: Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. However, existing methods generally l

safetyarxiv-cs-lg
21 Apr 2026
Safety

Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning

DGX agent

arXiv:2604.16702v1 Announce Type: new Abstract: Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestria

safetyarxiv-cs-ro
21 Apr 2026
Safety

Characterizing Model-Native Skills

DGX agent

arXiv:2604.17614v1 Announce Type: cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on

safetyarxiv-cs-cl
21 Apr 2026
Safety

Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation

DGX agent

arXiv:2604.17178v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) plays a critical role in mental health assistance by providing accessible psychological support in real-world appli

safetyarxiv-cs-cl
21 Apr 2026
Safety

Detecting Alarming Student Verbal Responses using Text and Audio Classifier

DGX agent

arXiv:2604.16717v1 Announce Type: new Abstract: This paper addresses a critical safety gap in the use Automated Verbal Response Scoring (AVRS). We present a novel hybrid framework for troubled student

safetyarxiv-cs-cl
21 Apr 2026
Local Ai

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks

DGX agent

arXiv:2604.18510v1 Announce Type: cross Abstract: Open-weight language models can be rendered unsafe through several distinct interventions, but the resulting models may differ substantially in capabi

local-aiarxiv-cs-cl
21 Apr 2026
Safety

Efficient Diffusion Models under Nonconvex Equality and Inequality constraints via Landing

DGX agent

arXiv:2604.17838v1 Announce Type: new Abstract: Generative modeling within constrained sets is essential for scientific and engineering applications involving physical, geometric, or safety requiremen

safetyarxiv-cs-lg
21 Apr 2026
Safety

Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models

DGX agent

arXiv:2604.16481v1 Announce Type: new Abstract: Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable

safetyarxiv-cs-cv
21 Apr 2026
Safety

HAVEN: Hierarchical Adversary-aware Visibility-Enabled Navigation with Cover Utilization using Deep Transformer Q-Networks

DGX agent

arXiv:2512.00592v2 Announce Type: replace Abstract: Autonomous navigation in partially observable environments requires agents to reason beyond immediate sensor input, exploit occlusion, and ensure sa

safetyarxiv-cs-ro
21 Apr 2026
Safety

Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization

DGX agent

arXiv:2604.17984v1 Announce Type: new Abstract: Uncertainty quantification is crucial in safety-critical systems, where decisions must be made under uncertainty. In particular, we consider the problem

safetyarxiv-cs-lg
21 Apr 2026
Safety

The Impact of Off-Policy Training Data on Probe Generalisation

DGX agent

arXiv:2511.17408v4 Announce Type: replace-cross Abstract: Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning b

safetyarxiv-cs-lg
21 Apr 2026
Safety

AI Agents and Hard Choices

DGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

safetyarxiv-cs-ai
20 Apr 2026
Safety

Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover

DGX agent

arXiv:2603.11331v2 Announce Type: replace-cross Abstract: Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. Empirically, we find that strong adversari

safetyarxiv-cs-ai
20 Apr 2026
Safety

Safe and Energy-Aware Multi-Robot Density Control via PDE-Constrained Optimization for Long-Duration Autonomy

DGX agent

arXiv:2604.15524v1 Announce Type: cross Abstract: This paper presents a novel density control framework for multi-robot systems with spatial safety and energy sustainability guarantees. Stochastic rob

safetyarxiv-cs-ro
20 Apr 2026
Safety

Safe Deep Reinforcement Learning for Building Heating Control and Demand-side Flexibility

DGX agent

arXiv:2604.16033v1 Announce Type: cross Abstract: Buildings account for approximately 40% of global energy consumption, and with the growing share of intermittent renewable energy sources, enabling de

safetyarxiv-cs-ai
20 Apr 2026
Safety

Context Over Content: Exposing Evaluation Faking in Automated Judges

DGX agent

arXiv:2604.15224v1 Announce Type: cross Abstract: The extit{LLM-as-a-judge} paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: th

safetyarxiv-cs-cl
17 Apr 2026
Safety

[Emerging Ideas] Artificial Tripartite Intelligence: A Bio-Inspired, Sensor-First Architecture for Physical AI

DGX agent

arXiv:2604.13959v1 Announce Type: new Abstract: As AI moves from data centers to robots and wearables, scaling ever-larger models becomes insufficient. Physical AI operates under tight latency, energy

safetyarxiv-cs-ai
17 Apr 2026
Safety

From Risk to Rescue: An Agentic Survival Analysis Framework for Liquidation Prevention

DGX agent

arXiv:2604.14583v1 Announce Type: new Abstract: Decentralized Finance (DeFi) lending protocols like Aave v3 rely on over-collateralization to secure loans, yet users frequently face liquidation due to

safetyarxiv-cs-lg
17 Apr 2026
Safety

Integrating Object Detection, LiDAR-Enhanced Depth Estimation, and Segmentation Models for Railway Environments

DGX agent

arXiv:2604.14781v1 Announce Type: new Abstract: Obstacle detection in railway environments is crucial for ensuring safety. However, very few studies address the problem using a complete, modular, and

safetyarxiv-cs-cv
17 Apr 2026
Safety

Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees

DGX agent

arXiv:2604.14243v1 Announce Type: new Abstract: Real-world decision-making systems operate in environments where state transitions depend not only on the agent's actions, but also on extbf{exogenous f

safetyarxiv-cs-lg
17 Apr 2026
Safety

The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious

DGX agent

arXiv:2604.14414v1 Announce Type: new Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However

safetyarxiv-cs-cl
17 Apr 2026
Safety

C^2T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination

DGX agent

arXiv:2604.13098v1 Announce Type: cross Abstract: State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (

safetyarxiv-cs-cv
16 Apr 2026
Safety

From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution

DGX agent

arXiv:2604.14053v1 Announce Type: new Abstract: Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves infe

safetyarxiv-cs-cl
16 Apr 2026
Safety

Learning Probabilistic Responsibility Allocations for Multi-Agent Interactions

DGX agent

arXiv:2604.13128v1 Announce Type: cross Abstract: Human behavior in interactive settings is shaped not only by individual objectives but also by shared constraints with others, such as safety. Underst

safetyarxiv-cs-lg
16 Apr 2026
Safety

Robust Energy-Aware Routing for Air-Ground Cooperative Multi-UAV Delivery in Wind-Uncertain Environments

DGX agent

arXiv:2604.13441v1 Announce Type: new Abstract: Ensuring energy feasibility under wind uncertainty is critical for the safety and reliability of UAV delivery missions. In realistic truck-drone logisti

safetyarxiv-cs-ro
16 Apr 2026
Safety

See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones

DGX agent

arXiv:2604.13292v1 Announce Type: new Abstract: Autonomous drone delivery systems are rapidly advancing, but ensuring safe and reliable package drop-offs remains highly challenging in cluttered urban

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

Synthesis and Deployment of Maximal Robust Control Barrier Functions through Adversarial Reinforcement Learning

DGX agent

arXiv:2604.13192v1 Announce Type: cross Abstract: Robust control barrier functions (CBFs) provide a principled mechanism for smooth safety enforcement under worst-case disturbances. However, existing

model-releasesarxiv-cs-ro
16 Apr 2026
Safety

ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs

DGX agent

arXiv:2512.09953v2 Announce Type: replace-cross Abstract: Machine unlearning aims to remove the influence of specific data points from a trained model to satisfy privacy, copyright, and safety require

safetyarxiv-cs-lg
16 Apr 2026
Safety

Fragile Reconstruction: Adversarial Vulnerability of Reconstruction-Based Detectors for Diffusion-Generated Images

DGX agent

arXiv:2604.12781v1 Announce Type: new Abstract: Recently, detecting AI-generated images produced by diffusion-based models has attracted increasing attention due to their potential threat to safety. A

safetyarxiv-cs-cv
15 Apr 2026
Safety

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software

DGX agent

arXiv:2604.12994v1 Announce Type: cross Abstract: Logical vulnerabilities in software stem from flaws in program logic rather than memory safety, which can lead to critical security failures. Although

safetyarxiv-cs-ai
15 Apr 2026
Safety

Parallax: Why AI Agents That Think Must Never Act

DGX agent

arXiv:2604.12986v1 Announce Type: cross Abstract: Autonomous AI agents are rapidly transitioning from experimental tools to operational infrastructure, with projections that 80% of enterprise applicat

safetyarxiv-cs-ai
15 Apr 2026
Safety

Skill-informed Data-driven Haptic Nudges for High-dimensional Human Motor Learning

DGX agent

arXiv:2603.12583v2 Announce Type: replace Abstract: In this work, we propose a data-driven framework to design optimal haptic nudge feedback leveraging the learner's estimated skill to address the cha

safetyarxiv-cs-ro
15 Apr 2026
Safety

A Mamba-Based Multimodal Network for Multiscale Blast-Induced Rapid Structural Damage Assessment

DGX agent

arXiv:2604.11709v1 Announce Type: new Abstract: Accurate and rapid structural damage assessment (SDA) is crucial for post-disaster management, helping responders prioritise resources, plan rescues, an

safetyarxiv-cs-ai
14 Apr 2026
Safety

AI Organizations are More Effective but Less Aligned than Individual Agents

DGX agent

arXiv:2604.10290v1 Announce Type: new Abstract: AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that

safetyarxiv-cs-ai
14 Apr 2026
Safety

Cross-Cultural Value Awareness in Large Vision-Language Models

DGX agent

arXiv:2604.09945v1 Announce Type: cross Abstract: The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity t

safetyarxiv-cs-ai
14 Apr 2026
Safety

ExpressMM: Expressive Mobile Manipulation Behaviors in Human-Robot Interactions

DGX agent

arXiv:2604.05320v2 Announce Type: replace Abstract: Mobile manipulators are increasingly deployed in human-centered environments to perform tasks. While completing such tasks, they should also be able

safetyarxiv-cs-ro
14 Apr 2026
Safety

Latent Structure of Affective Representations in Large Language Models

DGX agent

arXiv:2604.07382v2 Announce Type: replace-cross Abstract: The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implica

safetyarxiv-cs-ai
14 Apr 2026
Safety

Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning

DGX agent

arXiv:2604.11699v1 Announce Type: cross Abstract: This work aims to improve the generalization of logic-based legal reasoning systems by integrating recent advances in NLP with legal-domain adaptive f

safetyarxiv-cs-ai
14 Apr 2026
Safety

Rebooting Microreboot: Architectural Support for Safe, Parallel Recovery in Microservice Systems

DGX agent

arXiv:2604.09963v1 Announce Type: cross Abstract: Microreboot enables fast recovery by restarting only the failing component, but in modern microservices naive restarts are unsafe: dense dependencies

safetyarxiv-cs-ai
14 Apr 2026
Safety

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control

DGX agent

arXiv:2601.06540v2 Announce Type: replace-cross Abstract: This paper proposes a novel reinforcement learning framework, named Self-Organizing Dual-buffer Adaptive Clustering Experience Replay (SODACER

safetyarxiv-cs-ai
14 Apr 2026
Safety

From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

DGX agent

arXiv:2604.08591v1 Announce Type: cross Abstract: Hallucinations in large ASR models present a critical safety risk. In this work, we propose the extit{Spectral Sensitivity Theorem}, which predicts a

safetyarxiv-cs-ai
13 Apr 2026
Safety

Multimodal Anomaly Detection for Human-Robot Interaction

DGX agent

arXiv:2604.09326v1 Announce Type: cross Abstract: Ensuring safety and reliability in human-robot interaction (HRI) requires the timely detection of unexpected events that could lead to system failures

safetyarxiv-cs-cv
13 Apr 2026
Safety

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

DGX agent

arXiv:2604.08815v1 Announce Type: new Abstract: Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-re

safetyarxiv-cs-cv
13 Apr 2026
Safety

Before Humans Join the Team: Diagnosing Coordination Failures in Healthcare Robot Team Simulation

DGX agent

arXiv:2508.04691v2 Announce Type: replace-cross Abstract: As humans move toward collaborating with coordinated robot teams, understanding how these teams coordinate and fail is essential for building

safetyarxiv-cs-ai
10 Apr 2026
← Previous
1…2829303132…257
Next →