AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested

DGX agent

arXiv:2605.11496v1 Announce Type: cross Abstract: Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and

safetyarxiv-cs-lg
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness

DGX agent

arXiv:2503.16072v4 Announce Type: replace-cross Abstract: Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet mo

safetyarxiv-cs-cl
13 May 2026
Safety

VNDUQE: Information-Theoretic Novelty Detection using Deep Variational Information Bottleneck

DGX agent

arXiv:2605.11551v1 Announce Type: cross Abstract: Detecting out-of-distribution (OOD) samples is critical for safe deployment of neural networks in safety-critical applications. While maximum softmax

safetyarxiv-cs-cv
13 May 2026
Safety

Agent-Sentry: Bounding LLM Agents via Execution Provenance

DGX agent

arXiv:2603.22868v2 Announce Type: replace-cross Abstract: Agentic computing systems, while immensely capable, raise serious security, privacy, and safety concerns. A key issue is that the full set of

safetyarxiv-cs-ai
12 May 2026
Safety

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

DGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

safetyarxiv-cs-ai
12 May 2026
Safety

Consensus Sampling for Safer Generative AI

DGX agent

arXiv:2511.09493v2 Announce Type: replace Abstract: Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions

safetyarxiv-cs-ai
12 May 2026
Safety

Constraint-Aware Reinforcement Learning via Adaptive Action Scaling

DGX agent

arXiv:2510.11491v3 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) seeks to mitigate unsafe behaviors that arise from exploration during training by reducing constraint violati

safetyarxiv-cs-lg
12 May 2026
Safety

GNN for Structural Displacement Prediction

DGX agent

arXiv:2605.08303v1 Announce Type: cross Abstract: Accurate prediction of structural displacements under external loading is fundamental to structural health monitoring and seismic safety assessment. A

safetyarxiv-cs-ai
12 May 2026
Safety

How LLMs Are Persuaded: A Few Attention Heads, Rerouted

DGX agent

arXiv:2605.09314v1 Announce Type: new Abstract: Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly und

safetyarxiv-cs-ai
12 May 2026
Model Releases

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

DGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

model-releasesarxiv-cs-ai
12 May 2026
Safety

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

DGX agent

arXiv:2602.03677v2 Announce Type: replace Abstract: Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliab

safetyarxiv-cs-cl
12 May 2026
Model Releases

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

DGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

model-releasesarxiv-cs-cl
12 May 2026
Safety

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

DGX agent

arXiv:2605.05682v2 Announce Type: replace-cross Abstract: Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI

safetyarxiv-cs-ai
12 May 2026
Safety

Positive Alignment: Artificial Intelligence for Human Flourishing

DGX agent

arXiv:2605.10310v1 Announce Type: new Abstract: Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of ali

safetyarxiv-cs-ai
12 May 2026
Safety

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

DGX agent

arXiv:2605.10293v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide gua

safetyarxiv-cs-ai
12 May 2026
Safety

Safe Exploration for Nonlinear Processes Using Online Gaussian Process Learning

DGX agent

arXiv:2605.09772v1 Announce Type: cross Abstract: This paper proposes a safe data-driven control framework for nonlinear systems with partially known dynamics. The method ensures stability and constra

safetyarxiv-cs-ro
12 May 2026
Safety

UAV-Assisted Scan-to-Simulation for Landslides Using Physics-Informed Gaussian Splatting

DGX agent

arXiv:2605.10715v1 Announce Type: new Abstract: Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines t

safetyarxiv-cs-cv
12 May 2026
Safety

VISTA: A Generative Egocentric Video Framework for Daily Assistance

DGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

safetyarxiv-cs-cl
12 May 2026
Safety

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

DGX agent

arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne

safetyarxiv-cs-cv
12 May 2026
Safety

A Systematic Investigation of The RL-Jailbreaker in LLMs

DGX agent

arXiv:2605.07032v1 Announce Type: cross Abstract: The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversa

safetyarxiv-cs-ai
11 May 2026
Safety

Approximation-Free Differentiable Oblique Decision Trees

DGX agent

arXiv:2605.07837v1 Announce Type: cross Abstract: Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabu

safetyarxiv-cs-ai
11 May 2026
Safety

BEAVER: An Efficient Deterministic LLM Verifier

DGX agent

arXiv:2512.05439v2 Announce Type: replace Abstract: As large language models (LLMs) transition from research prototypes to production systems, practitioners often need reliable methods to verify model

safetyarxiv-cs-ai
11 May 2026
Safety

Beyond 'I cannot fulfill this request': Alleviating Rigid Rejection in LLMs via Label Enhancement

DGX agent

arXiv:2605.07883v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often l

safetyarxiv-cs-cl
11 May 2026
Safety

Fortifying Time Series: DTW-Certified Robust Anomaly Detection

DGX agent

arXiv:2605.07690v1 Announce Type: new Abstract: Time-series anomaly detection is critical for ensuring safety in high-stakes applications, where robustness is a fundamental requirement rather than a m

safetyarxiv-cs-lg
11 May 2026
Safety

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines

DGX agent

arXiv:2605.07062v1 Announce Type: cross Abstract: AI agents are assuming active roles in Continuous Integration and Continuous Deployment (CI/CD) workflows, yet the research community lacks a shared v

safetyarxiv-cs-ai
11 May 2026
Safety

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

DGX agent

arXiv:2605.06696v1 Announce Type: new Abstract: Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. Howev

safetyarxiv-cs-ai
11 May 2026
Safety

How Value Induction Reshapes LLM Behaviour

DGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

safetyarxiv-cs-cl
11 May 2026
Safety

Many-to-Many Multi-Agent Pickup and Delivery

DGX agent

arXiv:2605.07835v1 Announce Type: new Abstract: Multi-robot systems in automated warehouses must manage continuous streams of pickup-and-delivery tasks while ensuring efficiency and safety. Prior work

safetyarxiv-cs-ro
11 May 2026
Safety

PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection

DGX agent

arXiv:2509.26272v3 Announce Type: replace Abstract: The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the

safetyarxiv-cs-cv
11 May 2026
Safety

From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation

DGX agent

arXiv:2605.04327v1 Announce Type: new Abstract: We propose an architecture for integrating high-level, human-provided safety rules and operator-aligned semantic preferences into autonomous robot navig

safetyarxiv-cs-ro
7 May 2026
Safety

LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy

DGX agent

arXiv:2605.04295v1 Announce Type: new Abstract: LLMs' overconfidence, particularly when hallucinating, poses a significant challenge for the deployment of the models in safety-critical settings and ma

safetyarxiv-cs-lg
7 May 2026
Safety

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate

DGX agent

arXiv:2605.03409v1 Announce Type: new Abstract: We present Robust Agent Compensation (RAC), a log-based recovery paradigm (providing a safety net) implemented through an architectural extension that c

safetyarxiv-cs-ai
7 May 2026
Safety

A Three-Stage Offline SDRE-Based Control Framework for Human Motion Reproduction on a Suspended Bipedal Robot

DGX agent

arXiv:2506.04680v2 Announce Type: replace Abstract: During the development of wearable exoskeletons, evaluations involving human subjects pose inherent safety risks. Therefore, systematic testing is o

safetyarxiv-cs-ro
6 May 2026
Safety

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

DGX agent

arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing

safetyarxiv-cs-ai
6 May 2026
Safety

Efficient Temporal Datalog Materialisation for Composite Event Recognition

DGX agent

arXiv:2605.02488v1 Announce Type: new Abstract: Several applications demand the timely detection of critical situations, such as threats to safety and transparency, over high-velocity streams of symbo

safetyarxiv-cs-ai
6 May 2026
Safety

Human-in-the-Loop Uncertainty Analysis in Self-Adaptive Robots Using LLMs

DGX agent

arXiv:2605.02983v1 Announce Type: new Abstract: Self-adaptive robots operate in dynamic, unpredictable environments where unaddressed uncertainties can lead to safety violations and operational failur

safetyarxiv-cs-ro
6 May 2026
Safety

Instance-Level Costs for Nuanced Classifier Evaluation

DGX agent

arXiv:2605.03135v1 Announce Type: new Abstract: Standard classification treats all errors equally, but in content moderation, medical screening, and safety-critical applications, mistakes on clear-cut

safetyarxiv-cs-lg
6 May 2026
Safety

Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings

DGX agent

arXiv:2605.02908v1 Announce Type: new Abstract: Understanding how textual embeddings contribute to memorization in text-to-image diffusion models is crucial for both interpretability and safety. This

safetyarxiv-cs-cv
6 May 2026
Safety

Model Routing as a Trust Problem: Route Receipts for Adaptive AI Systems

DGX agent

arXiv:2605.01710v1 Announce Type: new Abstract: AI products often route requests through version aliases, service tiers, tool choices, regional endpoints, fallback rules, or safety handling before res

safetyarxiv-cs-ai
6 May 2026
Local Ai

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs

DGX agent

arXiv:2605.02946v1 Announce Type: new Abstract: Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly

local-aiarxiv-cs-lg
6 May 2026
Safety

Steerable Adversarial Scenario Generation through Test-Time Preference Alignment

DGX agent

arXiv:2509.20102v2 Announce Type: cross Abstract: Adversarial scenario generation is a cost-effective approach for safety assessment of autonomous driving systems. However, existing methods are often

safetyarxiv-cs-ro
6 May 2026
Safety

Zero-Shot Signal Temporal Logic Planning with Disjunctive Branch Selection in Dynamic Semantic Maps

DGX agent

arXiv:2605.01222v1 Announce Type: new Abstract: Signal Temporal Logic (STL) offers verifiable task specifications and is crucial for safety-critical control. Yet STL planning remains challenging: exac

safetyarxiv-cs-ai
6 May 2026
Safety

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

DGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

safetyarxiv-cs-cv
5 May 2026
Safety

Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure

DGX agent

arXiv:2605.00055v1 Announce Type: cross Abstract: We report a safety incident in a deployed multi-agent research system in which a primary AI agent installed 107 unauthorized software components, over

safetyarxiv-cs-ai
5 May 2026
Safety

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

DGX agent

arXiv:2605.01451v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly being integrated into high-stakes public safety systems, including emergency call triage and dispatch decision

safetyarxiv-cs-cl
5 May 2026
Safety

Causal Foundations of Collective Agency

DGX agent

arXiv:2605.00248v1 Announce Type: new Abstract: A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with c

safetyarxiv-cs-ai
5 May 2026
Safety

Combining Facial Videos and Biosignals for Stress Estimation During Driving

DGX agent

arXiv:2601.04376v3 Announce Type: replace Abstract: Reliable stress recognition is critical in applications such as medical monitoring and safety-critical systems, including real-world driving. While

safetyarxiv-cs-cv
5 May 2026
Safety

Edge Case Detection in Automated Driving: Methods, Challenges and Future Directions

DGX agent

arXiv:2410.08491v2 Announce Type: replace Abstract: Automated vehicles promise to enhance transportation safety and efficiency. However, ensuring their reliability in real-world conditions remains cha

safetyarxiv-cs-ro
5 May 2026
← Previous
1…2627282930…257
Next →