AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
19 May 2026

Transformer-Based MCS Prediction for 5G Multicast-Broadcast Services (MBS)

SafetyDGX agent

arXiv:2605.16735v1 Announce Type: cross Abstract: The deployment of 5G Multicast-Broadcast Services (MBS) is emerging as a critical technology for spectral-efficient UHD content delivery and serving a

Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2602.22801v2 Announce Type: replace-cross Abstract: Diffusion models have become a popular choice for decision-making tasks in robotics, and more recently, are also being considered for solving

What we find most useful about CNA is that the intervention is simple yet powerful. The steering is a multiplicative ablation on a sparse se…

SafetyDGX agent

What we find most useful about CNA is that the intervention is simple yet powerful. The steering is a multiplicative ablation on a sparse set of MLP neurons, which makes CNA a clean addition on top of

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
18 May 2026

AI Consciousness and Existential Risk

SafetyDGX agent

arXiv:2511.19115v2 Announce Type: replace Abstract: In AI, the existential risk denotes the hypothetical threat posed by an artificial system that would possess both the capability and the objective,

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.15687v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) may memorize sensitive cross-modal information during pretraining, making machine unlearning (MU) crucial. Ex

Constrained MPC-Based Motion Planning for Morphing Quadrotors in Ultra-Narrow Passages under Limited Perception

SafetyDGX agent

arXiv:2605.15999v1 Announce Type: new Abstract: This paper introduces a motion planning framework to plan morphology and trajectory for morphing quadrotors under extremely constrained environments. We

CTF4Nuclear: Common Task Framework for Nuclear Fission and Fusion Models

SafetyDGX agent

arXiv:2605.15549v1 Announce Type: cross Abstract: The demand for clean energy is ever increasing, with new nuclear technologies presenting a complementary solution to renewable energies. However, desi

Driving Through the Network: Performance and Workload Under Latency and Video Impairment

SafetyDGX agent

arXiv:2605.15952v1 Announce Type: cross Abstract: Teleoperation promises to extend the operational envelope of automated vehicles, yet it critically depends on network latency and video quality. We re

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems

SafetyDGX agent

arXiv:2605.16198v1 Announce Type: new Abstract: We examine one particular dimension of AI governance: how to monitor and audit AI-enabled products and services throughout the AI development lifecycle,

How Google Does It: Fleet-wide, large-scale A/B experimentation

SafetyDGX agent

When most people think of A/B experimentation, they think of button colors, landing page layouts, or checkout flows. At Google, many fundamental infrastructure improvements also need the rigor of A/B

Learning in Structured Stackelberg Games

SafetyDGX agent

arXiv:2504.09006v4 Announce Type: replace-cross Abstract: We initiate the study of structured Stackelberg games, a novel form of strategic interaction between a leader and a follower where contextual

OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation

SafetyDGX agent

arXiv:2605.15971v1 Announce Type: new Abstract: While reinforcement learning (RL) enables robots to acquire skills autonomously, its real-world deployment is severely limited by inefficient and unsafe

Quantum Artificial Intelligence for Mission-Critical Systems: Foundations, Architectural Elements, and Future Directions

SafetyDGX agent

arXiv:2511.09884v2 Announce Type: replace Abstract: Mission critical (MC) applications such as defense operations, energy management, cybersecurity, and aerospace control require reliable, determinist

SAFE Quantum Machine Learning with Variational Quantum Classifiers

SafetyDGX agent

arXiv:2605.16067v1 Announce Type: new Abstract: We propose a variational quantum classifier operating on high dimensional deep representations via amplitude encoding, stabilized by a learnable classic

Towards Trustworthy and Explainable AI for Perception Models: From Concept to Prototype Vehicle Deployment

SafetyDGX agent

arXiv:2605.16087v1 Announce Type: cross Abstract: Deep Neural Networks have become the dominant solution for Autonomous Driving perception, but their opacity conflicts with emerging Trustworthy AI gui

17 May 2026

What I am about to describe ain’t AGI; it’s a sign of a trillion dollar trainwreck. If I had told you in 2022 that the 2026 version of GPT (…

SafetyDGX agent

What I am about to describe ain’t AGI; it’s a sign of a trillion dollar trainwreck. If I had told you in 2022 that the 2026 version of GPT (which by the way would only be GPT 5.5 and not GPT-6 or 7 li

15 May 2026

A Regret Perspective on Online Multiple Testing

SafetyDGX agent

arXiv:2605.13916v1 Announce Type: cross Abstract: Online Multiple Testing (OMT), a fundamental pillar of sequential statistical inference, traditionally evaluates the False Discovery Rate (FDR) and st

Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model

SafetyDGX agent

arXiv:2605.14723v1 Announce Type: new Abstract: Sepsis management in the ICU requires sequential treatment decisions under rapidly evolving patient physiology. Although large language models (LLMs) en

Behavioral Data-Driven Optimal Trajectory Generation for Rotary Cranes

SafetyDGX agent

arXiv:2605.14944v1 Announce Type: new Abstract: With the growth of the construction industry and the global shortage of skilled labor, the automation of crane control has become increasingly important

DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2507.04049v4 Announce Type: replace Abstract: Most end-to-end autonomous driving methods rely on imitation learning from single expert demonstrations, often leading to conservative and homogeneo

Fair and Calibrated Toxicity Detection with Robust Training and Abstention

SafetyDGX agent

arXiv:2605.14074v1 Announce Type: new Abstract: Fairness in toxicity classification involves three integrated axes: ranking, calibration, and abstention. Training-time interventions and post-hoc safet

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2605.14201v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to bein

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs

SafetyDGX agent

arXiv:2605.15172v1 Announce Type: cross Abstract: Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safe

Novel Dynamic Batch-Sensitive Adam Optimiser for Vehicular Accident Injury Severity Prediction

SafetyDGX agent

arXiv:2605.15083v1 Announce Type: cross Abstract: The choice of optimiser is important in deep learning, as it strongly influences model efficiency and speed of convergence. However, many commonly use

Quantifying and Mitigating Premature Closure in Frontier LLMs

SafetyDGX agent

arXiv:2605.15000v1 Announce Type: cross Abstract: Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remai

Viverra: Text-to-Code with Guarantees

SafetyDGX agent

arXiv:2605.14972v1 Announce Type: cross Abstract: A fundamental limitation of Text-to-Code is that no guarantee can be obtained about the correctness of the generated code. Therefore, to ensure its co

14 May 2026

Asynchronous Reasoning: Training-Free Interactive Thinking LLMs

SafetyDGX agent

arXiv:2512.10931v3 Announce Type: replace Abstract: Many state-of-the-art LLMs are trained to think before giving their answer. Reasoning can greatly improve language model capabilities, but it also m

Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems

SafetyDGX agent

arXiv:2601.15161v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct

BEHAVE: A Hybrid AI Framework for Real-Time Modeling of Collective Human Dynamics

SafetyDGX agent

arXiv:2605.12730v1 Announce Type: new Abstract: Existing AI systems for modeling human behavior operate at the level of individuals or detect events after they occur. As a result, they systematically

DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense

SafetyDGX agent

arXiv:2605.13115v1 Announce Type: cross Abstract: Diffusion models depend on pseudo-random number generators (PRNGs) for latent noise sampling. We present DiffusionHijack, a supply-chain backdoor atta

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger sin…

SafetyDGX agent

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger single model get us there or a multi-agent system. The authors

Loiter UAV Reinsertion Guidance for Fixed-wing UAV Corridors

SafetyDGX agent

arXiv:2605.13822v1 Announce Type: new Abstract: This paper considers fixed-wing unmanned aerial vehicle (UAV) corridors comprising a main lane, a circular loiter lane for managing traffic congestion,

Neurosymbolic Auditing of Natural-Language Software Requirements

Model ReleasesDGX agent

arXiv:2605.13817v1 Announce Type: cross Abstract: Natural-language software requirements are often ambiguous, inconsistent, and underspecified; in safety-critical domains, these defects propagate into

No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

SafetyDGX agent

arXiv:2605.13044v1 Announce Type: cross Abstract: LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, b

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

SafetyDGX agent

arXiv:2605.12651v1 Announce Type: new Abstract: Runtime monitoring of autonomous systems traditionally relies on mapping continuous sensor observations to discrete logical propositions defined over lo

The Unified Autonomy Stack: Toward a Blueprint for Generalizable Robot Autonomy

SafetyDGX agent

arXiv:2605.12735v1 Announce Type: new Abstract: We introduce and open-source the Unified Autonomy Stack, a system-level solution that enables resilient autonomy across diverse aerial and ground robot

VERA-MH: Validation of Ethical and Responsible AI in Mental Health

SafetyDGX agent

arXiv:2605.13318v1 Announce Type: new Abstract: Chatbot usage has increased, including in fields for which they were never developed for--notably mental health support. To that end, we introduce Valid

13 May 2026

Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

SafetyDGX agent

arXiv:2605.11442v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have emerged as key intermediaries, orchestrating complex interactions between human users and a wide range of digit

Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing

SafetyDGX agent

arXiv:2505.05665v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated success in decision-making tasks including planning, control, and prediction, but thei

Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing

SafetyDGX agent

arXiv:2605.11202v1 Announce Type: cross Abstract: LLM inference and serving systems have become security-critical infrastructure; however, many of their most concerning failures arise from the serving

Leveraging RAG for Training-Free Alignment of LLMs

SafetyDGX agent

arXiv:2605.11217v1 Announce Type: new Abstract: Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to enabl

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

SafetyDGX agent

arXiv:2605.05630v2 Announce Type: replace Abstract: Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objec

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

SafetyDGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

The DAWN of World-Action Interactive Models

SafetyDGX agent

arXiv:2605.11550v1 Announce Type: new Abstract: A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action M

The new era of SaMD: Why cloud infrastructure is the foundation for digital health in 2026

SafetyDGX agent

In the healthcare and life sciences industries, speed saves lives, but meeting regulatory requirements and other administrative burdens often pumps the brakes for manufacturers of software as a medica

there are many agent use cases locked behind 'what if' fears because after we give a tool to an agent, we're relying on the prompt to limit …

SafetyDGX agent

there are many agent use cases locked behind 'what if' fears because after we give a tool to an agent, we're relying on the prompt to limit behavior tools like @denieddotdev allow teams to manage capa

12 May 2026

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

SafetyDGX agent

arXiv:2510.18184v3 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse aut

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

SafetyDGX agent

arXiv:2605.09902v1 Announce Type: new Abstract: Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risk

Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution

SafetyDGX agent

arXiv:2605.10223v1 Announce Type: new Abstract: Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk wr

Beyond Self-Play: Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation

SafetyDGX agent

arXiv:2605.09153v1 Announce Type: cross Abstract: Closed-loop traffic simulation requires agents that are both scalable and behaviorally realistic. Recent self-play reinforcement learning approaches d

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

SafetyDGX agent

arXiv:2605.10764v1 Announce Type: cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability,

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures

SafetyDGX agent

arXiv:2605.10770v1 Announce Type: new Abstract: Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains,

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

SafetyDGX agent

arXiv:2602.05946v3 Announce Type: replace Abstract: Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-pre

Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis

SafetyDGX agent

arXiv:2605.08139v1 Announce Type: cross Abstract: In modern distributed cloud environments, efficient resource allocation is required as traditional scaling mechanisms are often subject to cloud thras

MedFL-Stress: A Systematic Robustness Evaluation of Federated Brain Tumor Segmentation under Cross-Hospital MRI Appearance Shift

SafetyDGX agent

arXiv:2605.09025v1 Announce Type: new Abstract: Federated learning enables hospitals to collaboratively train segmentation models without sharing patient data. However, current evaluation protocols re

MoMo: Conditioned Contrastive Representation Learning for Preference-Modulated Planning

SafetyDGX agent

arXiv:2605.08512v1 Announce Type: new Abstract: Temporally contrastive representation learning induces a latent structure capable of reducing long-horizon planning to inference in a low-dimensional li

Muninn: Your Trajectory Diffusion Model But Faster

SafetyDGX agent

arXiv:2605.09999v1 Announce Type: new Abstract: Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohi

Mutual Information Optimal Density Control of Linear Systems and Generalized Schrodinger Bridges with Reference Refinement

SafetyDGX agent

arXiv:2605.09349v1 Announce Type: cross Abstract: We consider a mutual information (MI) regularized version of optimal density control of a discrete-time linear system. MI optimal control has been pro

Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

SafetyDGX agent

arXiv:2605.08778v1 Announce Type: new Abstract: Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based

← Previous
1…4445464748…240
Next →