AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
21 Apr 2026

Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation

SafetyDGX agent

arXiv:2604.17178v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) plays a critical role in mental health assistance by providing accessible psychological support in real-world appli

Detecting Alarming Student Verbal Responses using Text and Audio Classifier

SafetyDGX agent

arXiv:2604.16717v1 Announce Type: new Abstract: This paper addresses a critical safety gap in the use Automated Verbal Response Scoring (AVRS). We present a novel hybrid framework for troubled student

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks

Local AiDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.18510v1 Announce Type: cross Abstract: Open-weight language models can be rendered unsafe through several distinct interventions, but the resulting models may differ substantially in capabi

Efficient Diffusion Models under Nonconvex Equality and Inequality constraints via Landing

SafetyDGX agent

arXiv:2604.17838v1 Announce Type: new Abstract: Generative modeling within constrained sets is essential for scientific and engineering applications involving physical, geometric, or safety requiremen

Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models

SafetyDGX agent

arXiv:2604.16481v1 Announce Type: new Abstract: Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable

HAVEN: Hierarchical Adversary-aware Visibility-Enabled Navigation with Cover Utilization using Deep Transformer Q-Networks

SafetyDGX agent

arXiv:2512.00592v2 Announce Type: replace Abstract: Autonomous navigation in partially observable environments requires agents to reason beyond immediate sensor input, exploit occlusion, and ensure sa

Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization

SafetyDGX agent

arXiv:2604.17984v1 Announce Type: new Abstract: Uncertainty quantification is crucial in safety-critical systems, where decisions must be made under uncertainty. In particular, we consider the problem

The Impact of Off-Policy Training Data on Probe Generalisation

SafetyDGX agent

arXiv:2511.17408v4 Announce Type: replace-cross Abstract: Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning b

They just indicted the Southern Poverty Law Center on 11 counts of fraud. Wow! The SPLC worked with Facebook, Old Twitter and Youtube on the…

SafetyDGX agent

They just indicted the Southern Poverty Law Center on 11 counts of fraud. Wow! The SPLC worked with Facebook, Old Twitter and Youtube on their safety teams, sometimes directly advising/flagging what c

20 Apr 2026

AI Agents and Hard Choices

SafetyDGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover

SafetyDGX agent

arXiv:2603.11331v2 Announce Type: replace-cross Abstract: Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. Empirically, we find that strong adversari

Safe and Energy-Aware Multi-Robot Density Control via PDE-Constrained Optimization for Long-Duration Autonomy

SafetyDGX agent

arXiv:2604.15524v1 Announce Type: cross Abstract: This paper presents a novel density control framework for multi-robot systems with spatial safety and energy sustainability guarantees. Stochastic rob

Safe Deep Reinforcement Learning for Building Heating Control and Demand-side Flexibility

SafetyDGX agent

arXiv:2604.16033v1 Announce Type: cross Abstract: Buildings account for approximately 40% of global energy consumption, and with the growing share of intermittent renewable energy sources, enabling de

18 Apr 2026

Dario is not as different from Sam as people seem to think.

SafetyDGX agent

Gary Marcus argues that Dario Amodei (Anthropic CEO) and Sam Altman (OpenAI CEO) share more similarities in their approaches to AI development and safety than commonly perceived, despite their public

17 Apr 2026

Context Over Content: Exposing Evaluation Faking in Automated Judges

SafetyDGX agent

arXiv:2604.15224v1 Announce Type: cross Abstract: The extit{LLM-as-a-judge} paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: th

[Emerging Ideas] Artificial Tripartite Intelligence: A Bio-Inspired, Sensor-First Architecture for Physical AI

SafetyDGX agent

arXiv:2604.13959v1 Announce Type: new Abstract: As AI moves from data centers to robots and wearables, scaling ever-larger models becomes insufficient. Physical AI operates under tight latency, energy

From Risk to Rescue: An Agentic Survival Analysis Framework for Liquidation Prevention

SafetyDGX agent

arXiv:2604.14583v1 Announce Type: new Abstract: Decentralized Finance (DeFi) lending protocols like Aave v3 rely on over-collateralization to secure loans, yet users frequently face liquidation due to

Integrating Object Detection, LiDAR-Enhanced Depth Estimation, and Segmentation Models for Railway Environments

SafetyDGX agent

arXiv:2604.14781v1 Announce Type: new Abstract: Obstacle detection in railway environments is crucial for ensuring safety. However, very few studies address the problem using a complete, modular, and

Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees

SafetyDGX agent

arXiv:2604.14243v1 Announce Type: new Abstract: Real-world decision-making systems operate in environments where state transitions depend not only on the agent's actions, but also on extbf{exogenous f

The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious

SafetyDGX agent

arXiv:2604.14414v1 Announce Type: new Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However

16 Apr 2026

C^2T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination

SafetyDGX agent

arXiv:2604.13098v1 Announce Type: cross Abstract: State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (

From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution

SafetyDGX agent

arXiv:2604.14053v1 Announce Type: new Abstract: Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves infe

Learning Probabilistic Responsibility Allocations for Multi-Agent Interactions

SafetyDGX agent

arXiv:2604.13128v1 Announce Type: cross Abstract: Human behavior in interactive settings is shaped not only by individual objectives but also by shared constraints with others, such as safety. Underst

Robust Energy-Aware Routing for Air-Ground Cooperative Multi-UAV Delivery in Wind-Uncertain Environments

SafetyDGX agent

arXiv:2604.13441v1 Announce Type: new Abstract: Ensuring energy feasibility under wind uncertainty is critical for the safety and reliability of UAV delivery missions. In realistic truck-drone logisti

See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones

SafetyDGX agent

arXiv:2604.13292v1 Announce Type: new Abstract: Autonomous drone delivery systems are rapidly advancing, but ensuring safe and reliable package drop-offs remains highly challenging in cluttered urban

Synthesis and Deployment of Maximal Robust Control Barrier Functions through Adversarial Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.13192v1 Announce Type: cross Abstract: Robust control barrier functions (CBFs) provide a principled mechanism for smooth safety enforcement under worst-case disturbances. However, existing

ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs

SafetyDGX agent

arXiv:2512.09953v2 Announce Type: replace-cross Abstract: Machine unlearning aims to remove the influence of specific data points from a trained model to satisfy privacy, copyright, and safety require

15 Apr 2026

Fragile Reconstruction: Adversarial Vulnerability of Reconstruction-Based Detectors for Diffusion-Generated Images

SafetyDGX agent

arXiv:2604.12781v1 Announce Type: new Abstract: Recently, detecting AI-generated images produced by diffusion-based models has attracted increasing attention due to their potential threat to safety. A

FSD v14.3.1 review after 10 drives and many hours, here are my thoughts (it’s a great one) - v14.3.1 feels like a big jump even compared to …

SafetyDGX agent

FSD v14.3.1 review after 10 drives and many hours, here are my thoughts (it’s a great one) - v14.3.1 feels like a big jump even compared to FSD v14.3, even for “just” a point release build. Everything

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software

SafetyDGX agent

arXiv:2604.12994v1 Announce Type: cross Abstract: Logical vulnerabilities in software stem from flaws in program logic rather than memory safety, which can lead to critical security failures. Although

Parallax: Why AI Agents That Think Must Never Act

SafetyDGX agent

arXiv:2604.12986v1 Announce Type: cross Abstract: Autonomous AI agents are rapidly transitioning from experimental tools to operational infrastructure, with projections that 80% of enterprise applicat

Skill-informed Data-driven Haptic Nudges for High-dimensional Human Motor Learning

SafetyDGX agent

arXiv:2603.12583v2 Announce Type: replace Abstract: In this work, we propose a data-driven framework to design optimal haptic nudge feedback leveraging the learner's estimated skill to address the cha

14 Apr 2026

A Mamba-Based Multimodal Network for Multiscale Blast-Induced Rapid Structural Damage Assessment

SafetyDGX agent

arXiv:2604.11709v1 Announce Type: new Abstract: Accurate and rapid structural damage assessment (SDA) is crucial for post-disaster management, helping responders prioritise resources, plan rescues, an

AI Organizations are More Effective but Less Aligned than Individual Agents

SafetyDGX agent

arXiv:2604.10290v1 Announce Type: new Abstract: AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that

Cross-Cultural Value Awareness in Large Vision-Language Models

SafetyDGX agent

arXiv:2604.09945v1 Announce Type: cross Abstract: The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity t

ExpressMM: Expressive Mobile Manipulation Behaviors in Human-Robot Interactions

SafetyDGX agent

arXiv:2604.05320v2 Announce Type: replace Abstract: Mobile manipulators are increasingly deployed in human-centered environments to perform tasks. While completing such tasks, they should also be able

Latent Structure of Affective Representations in Large Language Models

SafetyDGX agent

arXiv:2604.07382v2 Announce Type: replace-cross Abstract: The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implica

Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning

SafetyDGX agent

arXiv:2604.11699v1 Announce Type: cross Abstract: This work aims to improve the generalization of logic-based legal reasoning systems by integrating recent advances in NLP with legal-domain adaptive f

Rebooting Microreboot: Architectural Support for Safe, Parallel Recovery in Microservice Systems

SafetyDGX agent

arXiv:2604.09963v1 Announce Type: cross Abstract: Microreboot enables fast recovery by restarting only the failing component, but in modern microservices naive restarts are unsafe: dense dependencies

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control

SafetyDGX agent

arXiv:2601.06540v2 Announce Type: replace-cross Abstract: This paper proposes a novel reinforcement learning framework, named Self-Organizing Dual-buffer Adaptive Clustering Experience Replay (SODACER

13 Apr 2026

From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

SafetyDGX agent

arXiv:2604.08591v1 Announce Type: cross Abstract: Hallucinations in large ASR models present a critical safety risk. In this work, we propose the extit{Spectral Sensitivity Theorem}, which predicts a

Multimodal Anomaly Detection for Human-Robot Interaction

SafetyDGX agent

arXiv:2604.09326v1 Announce Type: cross Abstract: Ensuring safety and reliability in human-robot interaction (HRI) requires the timely detection of unexpected events that could lead to system failures

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

SafetyDGX agent

arXiv:2604.08815v1 Announce Type: new Abstract: Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-re

10 Apr 2026

Before Humans Join the Team: Diagnosing Coordination Failures in Healthcare Robot Team Simulation

SafetyDGX agent

arXiv:2508.04691v2 Announce Type: replace-cross Abstract: As humans move toward collaborating with coordinated robot teams, understanding how these teams coordinate and fail is essential for building

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

SafetyDGX agent

arXiv:2604.06210v2 Announce Type: cross Abstract: As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing benchmarks fac

EMMa: End-Effector Stability-Oriented Mobile Manipulation for Tracked Rescue Robots

SafetyDGX agent

arXiv:2604.08292v1 Announce Type: new Abstract: The autonomous operation of tracked mobile manipulators in rescue missions requires not only ensuring the reachability and safety of robot motion but al

Evaluation as Evolution: Transforming Adversarial Diffusion into Closed-Loop Curricula for Autonomous Vehicles

SafetyDGX agent

arXiv:2604.07378v1 Announce Type: new Abstract: Autonomous vehicles in interactive traffic environments are often limited by the scarcity of safety-critical tail events in static datasets, which biase

Event-Centric World Modeling with Memory-Augmented Retrieval for Embodied Decision-Making

SafetyDGX agent

arXiv:2604.07392v1 Announce Type: cross Abstract: Autonomous agents operating in dynamic and safety-critical environments require decision-making frameworks that are both computationally efficient and

Event-Level Detection of Surgical Instrument Handovers in Videos with Interpretable Vision Models

SafetyDGX agent

arXiv:2604.07577v1 Announce Type: new Abstract: Reliable monitoring of surgical instrument exchanges is essential for maintaining procedural efficiency and patient safety in the operating room. Automa

Responsible and safe use of AI

SafetyDGX agent

The OpenAI Academy page on 'Responsible and Safe Use of AI' is a guidance resource focused on best practices for using ChatGPT responsibly in professional and personal settings. It emphasizes that ...

9 Apr 2026

Read The Story of Civilization by Durant

SafetyDGX agent

Read The Story of Civilization by Durant “But the curse of every ancient civilization was that its men in the end became unable to fight. Materialism, luxury, safety, even sometimes an almost modern s

8 Apr 2026

link to @HeidyKhlaaf’s sharp analysis:

SafetyDGX agent

link to @HeidyKhlaaf’s sharp analysis: As someone who has audited dozens of safety-critical systems, built static analysis tools, and used most formal verification and security tools, here are some re

13 Aug 2026

Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis

SafetyDGX agent

arXiv:2608.11480v1 Announce Type: cross Abstract: Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment

SafetyDGX agent

arXiv:2608.12198v1 Announce Type: cross Abstract: Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of au

12 Aug 2026

ELMER: Evolutionary Language Model that Explores and Refines

SafetyDGX agent

arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

SafetyDGX agent

arXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial r

How to Verify Consistency of Probabilistic Claims

SafetyDGX agent

arXiv:2608.11181v1 Announce Type: cross Abstract: When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial t

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

SafetyDGX agent

arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain chall

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

SafetyDGX agent

arXiv:2608.10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in signi

Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data Advantage

SafetyDGX agent

arXiv:2608.10590v1 Announce Type: new Abstract: Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on

← Previous
1…2526272829…240
Next →