AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

DGX agent

arXiv:2605.30273v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and s

safetyarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving

DGX agent

arXiv:2605.27763v1 Announce Type: new Abstract: Safety evaluations of language models often treat serving configuration as fixed background infrastructure, but batch condition is an untested treatment

local-aiarxiv-cs-lg
28 May 2026
Safety

Chance-Constrained MPPI under State and Dynamic Object Prediction Uncertainty and the Evaluation of Collision Risk Calibration

DGX agent

arXiv:2605.28330v1 Announce Type: new Abstract: Chance-constrained Model Predictive Path Integral (MPPI) control is increasingly adopted for navigation in dynamic environments to explicitly bound coll

safetyarxiv-cs-ro
28 May 2026
Safety

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

DGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

safetyarxiv-cs-ai
28 May 2026
Safety

Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?

DGX agent

arXiv:2605.27494v1 Announce Type: cross Abstract: Modern retrieval-augmented generation(RAG) deployments increasingly rely on caching to reduce token cost and time-to-first-token(TTFT). Prefix-level K

safetyarxiv-cs-ai
28 May 2026
Safety

LACUNA: Safe Agents as Recursive Program Holes

DGX agent

arXiv:2605.28617v1 Announce Type: new Abstract: LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime o

safetyarxiv-cs-ai
28 May 2026
Model Releases

Models That Know How Evaluations Are Designed Score Safer

DGX agent

arXiv:2605.28591v1 Announce Type: cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified tes

model-releasesarxiv-cs-ai
28 May 2026
Safety

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

DGX agent

arXiv:2605.06213v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor e

safetyarxiv-cs-ai
27 May 2026
Safety

Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial

DGX agent

arXiv:2605.26577v1 Announce Type: cross Abstract: Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However,

safetyarxiv-cs-ai
27 May 2026
Safety

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

DGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

safetyarxiv-cs-ai
27 May 2026
Safety

Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation

DGX agent

arXiv:2603.25415v2 Announce Type: replace Abstract: Semantic world models enable embodied agents to reason about objects, relations, and spatial context beyond purely geometric representations. In Org

safetyarxiv-cs-ai
27 May 2026
Safety

Provably Safe Motion Planning Under Unknown Disturbances

DGX agent

arXiv:2605.26625v1 Announce Type: new Abstract: We present a provably safe sampling-based motion planning algorithm for robotic systems affected by random disturbances of unknown distribution. We cons

safetyarxiv-cs-ro
27 May 2026
Safety

Stochastic Decision Horizons for Constrained Reinforcement Learning

DGX agent

arXiv:2602.04599v2 Announce Type: replace Abstract: We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint sat

safetyarxiv-cs-lg
27 May 2026
Safety

Auditing medical multi-agent AI reveals risks of false consensus

DGX agent

arXiv:2510.10185v2 Announce Type: replace-cross Abstract: Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through sp

safetyarxiv-cs-ai
26 May 2026
Safety

Energy Shields for Fairness

DGX agent

arXiv:2605.24926v1 Announce Type: new Abstract: Runtime fairness is not a one-time constraint but a dynamic property evaluated over a sequence of decisions. To ensure fairness at runtime, it is necess

safetyarxiv-cs-ai
26 May 2026
Safety

First, do no harm: Breaking suicidogenic echo chambers in media recommendation

DGX agent

arXiv:2605.25258v1 Announce Type: cross Abstract: Recommender systems generally optimises user engagement, but this approach is dangerous in mental health contexts. When vulnerable users show signs of

safetyarxiv-cs-ai
26 May 2026
Safety

On the Stability and Realizability of Recurrent Polynomial Surrogate Ternary Logic Gate Networks

DGX agent

arXiv:2605.24649v1 Announce Type: cross Abstract: Recurrent Neural Networks (RNNs) can learn to predict Signal Temporal Logic (STL) verdicts online from partial trajectories, but deploying them as run

safetyarxiv-cs-ai
26 May 2026
Safety

Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving

DGX agent

arXiv:2605.24004v1 Announce Type: new Abstract: Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic

safetyarxiv-cs-ai
26 May 2026
Model Releases

RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection

DGX agent

arXiv:2605.25095v1 Announce Type: new Abstract: Autonomous driving stacks must pick one trajectory from a multi-modal candidate set; choosing by model confidence ignores safety, traffic-law, and comfo

model-releasesarxiv-cs-ai
26 May 2026
Safety

Referential Security as a New Paradigm for AI Evaluations

DGX agent

arXiv:2605.25673v1 Announce Type: cross Abstract: Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact

safetyarxiv-cs-ai
26 May 2026
Safety

SEIDM: A Safe and Efficient Intelligent Driver Model for Autonomous Driving Behavior

DGX agent

arXiv:2605.23915v1 Announce Type: cross Abstract: The Intelligent Driver Model (IDM) is a cornerstone of Adaptive Cruise Control (ACC), valued for its interpretable parameters and effectiveness in car

safetyarxiv-cs-ro
26 May 2026
Safety

RAG-Pull: Turning Retrieval into a Code-Injection Channel via Invisible Unicode Perturbations

DGX agent

arXiv:2510.11195v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) increases the reliability and trustworthiness of the LLM response and reduces hallucination by eliminatin

safetyarxiv-cs-ai
25 May 2026
Safety

Branch-Stochastic Model Predictive Control for Motion Planning under Multi-Modal Uncertainty with Scenario Clustering

DGX agent

arXiv:2605.22600v1 Announce Type: new Abstract: Motion planning for autonomous driving must account for multi-modal uncertainty in both the intentions and trajectories of surrounding vehicles. Handlin

safetyarxiv-cs-ro
22 May 2026
Safety

Learning to Evolve: Multi-modal Interactive Fields for Robust Humanoid Navigation in Dynamic Environments

DGX agent

arXiv:2605.21935v1 Announce Type: new Abstract: Safe manipulation-oriented navigation for humanoid robots requires scene memory that remains reliable under locomotion-induced perceptual distortion, en

safetyarxiv-cs-ro
22 May 2026
Safety

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI

DGX agent

arXiv:2507.05660v3 Announce Type: replace-cross Abstract: Customizing Large Language Models (LLMs) on untrusted datasets poses severe risks of injecting toxic behaviors. In this work, we introduce Opt

safetyarxiv-cs-cl
22 May 2026
Safety

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.22748v1 Announce Type: new Abstract: Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This f

safetyarxiv-cs-ro
22 May 2026
Safety

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

DGX agent

arXiv:2605.21139v1 Announce Type: new Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement

safetyarxiv-cs-cv
21 May 2026
Safety

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

DGX agent

arXiv:2605.20654v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circ

safetyarxiv-cs-lg
21 May 2026
Safety

Neural Configuration-Space Barriers for Manipulation Planning and Control

DGX agent

arXiv:2503.04929v3 Announce Type: replace-cross Abstract: Planning and control for high-dimensional robot manipulators in cluttered dynamic environments require computational efficiency and robust saf

safetyarxiv-cs-lg
20 May 2026
Safety

Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks

DGX agent

arXiv:2605.18988v1 Announce Type: cross Abstract: The expansion of Multimodal Large Language Models (MLLMs) and their integration into autonomous agentic workflows has introduced a non-stationary atta

safetyarxiv-cs-ai
20 May 2026
Safety

Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings

DGX agent

arXiv:2605.16993v1 Announce Type: cross Abstract: Current clinical artificial intelligence (AI) systems are evaluated almost exclusively on clean, standardised, English-language inputs, conditions tha

safetyarxiv-cs-ai
19 May 2026
Safety

Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression

DGX agent

arXiv:2605.17304v1 Announce Type: cross Abstract: LLM context is not just tokens; it is a set of commitments. Long-running conversations accumulate goals, constraints, decisions, preferences, tool res

safetyarxiv-cs-cl
19 May 2026
Safety

Contrastive-SDXL: Annotation-Preserving Night-Time Augmentation for Pedestrian Detection

DGX agent

arXiv:2605.16406v1 Announce Type: new Abstract: Night-time pedestrian detection remains challenging because labelled night-time data are limited and large illumination differences make daytime-only tr

safetyarxiv-cs-cv
19 May 2026
Safety

Pedestrian-Aware LLM-Driven Behavioral Planning for Autonomous Vehicles

DGX agent

arXiv:2605.16858v1 Announce Type: cross Abstract: Autonomous Vehicles (AVs) must make reliable decisions in dense urban environments where pedestrian behavior is variable, sometimes abnormal, and ofte

safetyarxiv-cs-ai
19 May 2026
Safety

Prediction of Challenging Behaviors Associated with Profound Autism in a Classroom Setting Using Wearable Sensors

DGX agent

arXiv:2605.17618v1 Announce Type: new Abstract: Autism Spectrum Disorder (ASD) is characterized by challenges with social interaction and communication and by restricted or repetitive patterns of thou

safetyarxiv-cs-ai
19 May 2026
Safety

SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training

DGX agent

arXiv:2605.18719v1 Announce Type: new Abstract: Diffusion models have been widely studied for removing unsafe content learned during pre-training. Existing methods require expensive supervised data, e

safetyarxiv-cs-cv
19 May 2026
Safety

State Contamination in Memory-Augmented LLM Agents

DGX agent

arXiv:2605.16746v1 Announce Type: new Abstract: LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon inter

safetyarxiv-cs-ai
19 May 2026
Local Ai

EVA: Editing for Versatile Alignment against Jailbreaks

DGX agent

arXiv:2605.14750v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks

local-aiarxiv-cs-ai
15 May 2026
Safety

Synthesizing POMDP Policies: Sampling Meets Model-checking via Learning

DGX agent

arXiv:2605.14440v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are the standard framework for decision-making under uncertainty. While sampling-based methods s

safetyarxiv-cs-ai
15 May 2026
Safety

AgenticAITA: A Proof-Of-Concept About Deliberative Multi-Agent Reasoning for Autonomous Trading Systems

DGX agent

arXiv:2605.12532v1 Announce Type: cross Abstract: Conventional algorithmic trading systems are grounded in deterministic heuristics or offline-trained statistical models that cannot adapt to the seman

safetyarxiv-cs-ai
14 May 2026
Safety

Automated alignment is harder than you think

DGX agent

arXiv:2605.06390v2 Announce Type: replace Abstract: A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as c

safetyarxiv-cs-ai
14 May 2026
Safety

Beyond VMAF: Towards Application-Specific Metrics for Teleoperation Video

DGX agent

arXiv:2605.13525v1 Announce Type: cross Abstract: Automated driving has made remarkable progress, yet situations still arise where human intervention is necessary. Teleoperation provides a scalable so

safetyarxiv-cs-ro
14 May 2026
Safety

Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing

DGX agent

arXiv:2605.12876v1 Announce Type: new Abstract: Randomized smoothing provides strong, model-agnostic robustness certificates, but existing guarantees are limited to single modalities, treating continu

safetyarxiv-cs-lg
14 May 2026
Safety

History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions

DGX agent

arXiv:2605.13825v1 Announce Type: new Abstract: Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different mod

safetyarxiv-cs-ai
14 May 2026
Safety

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

DGX agent

arXiv:2605.12561v1 Announce Type: new Abstract: Safe reinforcement learning (RL) typically asks extit{what} an agent should do. We ask extit{when} it needs to act, and show that a single policy can jo

safetyarxiv-cs-lg
14 May 2026
Safety

Certified Gradient-Based Contact-Rich Manipulation via Smoothing-Error Reachable Tubes

DGX agent

arXiv:2602.09368v2 Announce Type: replace Abstract: Gradient-based methods can efficiently optimize controllers by leveraging differentiable simulation and physical priors. However, contact-rich manip

safetyarxiv-cs-ro
13 May 2026
Safety

Generative AI for Visualizing Highway Construction Hazards Through Synthetic Images and Temporal Sequences

DGX agent

arXiv:2605.11276v1 Announce Type: new Abstract: Highway construction workers face a high risk of serious injury or death. Image-based training materials depicting hazardous scenarios are essential for

safetyarxiv-cs-cv
13 May 2026
Safety

Robust Policy Optimization to Prevent Catastrophic Forgetting

DGX agent

arXiv:2602.08813v2 Announce Type: replace Abstract: Large language models are commonly trained through multi-stage post-training: first via RLHF, then fine-tuned for other downstream objectives. Yet e

safetyarxiv-cs-lg
13 May 2026
← Previous
1…3233343536…257
Next →