AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Safety

Prototype Fusion: A Training-Free Multi-Layer Approach to OOD Detection

DGX agent

arXiv:2603.23677v2 Announce Type: replace Abstract: Deep learning models are increasingly deployed in safety-critical applications, where reliable out-of-distribution (OOD) detection is essential to e

safetyarxiv-cs-cv
13 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested

DGX agent

arXiv:2605.11496v1 Announce Type: cross Abstract: Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and

safetyarxiv-cs-lg
13 May 2026
Safety

Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness

DGX agent

arXiv:2503.16072v4 Announce Type: replace-cross Abstract: Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet mo

safetyarxiv-cs-cl
13 May 2026
Safety

VNDUQE: Information-Theoretic Novelty Detection using Deep Variational Information Bottleneck

DGX agent

arXiv:2605.11551v1 Announce Type: cross Abstract: Detecting out-of-distribution (OOD) samples is critical for safe deployment of neural networks in safety-critical applications. While maximum softmax

safetyarxiv-cs-cv
13 May 2026
Safety

Agent-Sentry: Bounding LLM Agents via Execution Provenance

DGX agent

arXiv:2603.22868v2 Announce Type: replace-cross Abstract: Agentic computing systems, while immensely capable, raise serious security, privacy, and safety concerns. A key issue is that the full set of

safetyarxiv-cs-ai
12 May 2026
Safety

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

DGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

safetyarxiv-cs-ai
12 May 2026
Safety

Consensus Sampling for Safer Generative AI

DGX agent

arXiv:2511.09493v2 Announce Type: replace Abstract: Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions

safetyarxiv-cs-ai
12 May 2026
Safety

Constraint-Aware Reinforcement Learning via Adaptive Action Scaling

DGX agent

arXiv:2510.11491v3 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) seeks to mitigate unsafe behaviors that arise from exploration during training by reducing constraint violati

safetyarxiv-cs-lg
12 May 2026
Safety

GNN for Structural Displacement Prediction

DGX agent

arXiv:2605.08303v1 Announce Type: cross Abstract: Accurate prediction of structural displacements under external loading is fundamental to structural health monitoring and seismic safety assessment. A

safetyarxiv-cs-ai
12 May 2026
Safety

How LLMs Are Persuaded: A Few Attention Heads, Rerouted

DGX agent

arXiv:2605.09314v1 Announce Type: new Abstract: Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly und

safetyarxiv-cs-ai
12 May 2026
Model Releases

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

DGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

model-releasesarxiv-cs-ai
12 May 2026
Safety

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

DGX agent

arXiv:2602.03677v2 Announce Type: replace Abstract: Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliab

safetyarxiv-cs-cl
12 May 2026
Model Releases

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

DGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

model-releasesarxiv-cs-cl
12 May 2026
Safety

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

DGX agent

arXiv:2605.05682v2 Announce Type: replace-cross Abstract: Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI

safetyarxiv-cs-ai
12 May 2026
Safety

Positive Alignment: Artificial Intelligence for Human Flourishing

DGX agent

arXiv:2605.10310v1 Announce Type: new Abstract: Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of ali

safetyarxiv-cs-ai
12 May 2026
Safety

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

DGX agent

arXiv:2605.10293v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide gua

safetyarxiv-cs-ai
12 May 2026
Safety

Safe Exploration for Nonlinear Processes Using Online Gaussian Process Learning

DGX agent

arXiv:2605.09772v1 Announce Type: cross Abstract: This paper proposes a safe data-driven control framework for nonlinear systems with partially known dynamics. The method ensures stability and constra

safetyarxiv-cs-ro
12 May 2026
Safety

UAV-Assisted Scan-to-Simulation for Landslides Using Physics-Informed Gaussian Splatting

DGX agent

arXiv:2605.10715v1 Announce Type: new Abstract: Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines t

safetyarxiv-cs-cv
12 May 2026
Safety

VISTA: A Generative Egocentric Video Framework for Daily Assistance

DGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

safetyarxiv-cs-cl
12 May 2026
Safety

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

DGX agent

arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne

safetyarxiv-cs-cv
12 May 2026
Safety

A Systematic Investigation of The RL-Jailbreaker in LLMs

DGX agent

arXiv:2605.07032v1 Announce Type: cross Abstract: The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversa

safetyarxiv-cs-ai
11 May 2026
Safety

Approximation-Free Differentiable Oblique Decision Trees

DGX agent

arXiv:2605.07837v1 Announce Type: cross Abstract: Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabu

safetyarxiv-cs-ai
11 May 2026
Safety

BEAVER: An Efficient Deterministic LLM Verifier

DGX agent

arXiv:2512.05439v2 Announce Type: replace Abstract: As large language models (LLMs) transition from research prototypes to production systems, practitioners often need reliable methods to verify model

safetyarxiv-cs-ai
11 May 2026
Safety

Beyond 'I cannot fulfill this request': Alleviating Rigid Rejection in LLMs via Label Enhancement

DGX agent

arXiv:2605.07883v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often l

safetyarxiv-cs-cl
11 May 2026
Safety

Fortifying Time Series: DTW-Certified Robust Anomaly Detection

DGX agent

arXiv:2605.07690v1 Announce Type: new Abstract: Time-series anomaly detection is critical for ensuring safety in high-stakes applications, where robustness is a fundamental requirement rather than a m

safetyarxiv-cs-lg
11 May 2026
Safety

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines

DGX agent

arXiv:2605.07062v1 Announce Type: cross Abstract: AI agents are assuming active roles in Continuous Integration and Continuous Deployment (CI/CD) workflows, yet the research community lacks a shared v

safetyarxiv-cs-ai
11 May 2026
Safety

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

DGX agent

arXiv:2605.06696v1 Announce Type: new Abstract: Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. Howev

safetyarxiv-cs-ai
11 May 2026
Safety

How Value Induction Reshapes LLM Behaviour

DGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

safetyarxiv-cs-cl
11 May 2026
Safety

Many-to-Many Multi-Agent Pickup and Delivery

DGX agent

arXiv:2605.07835v1 Announce Type: new Abstract: Multi-robot systems in automated warehouses must manage continuous streams of pickup-and-delivery tasks while ensuring efficiency and safety. Prior work

safetyarxiv-cs-ro
11 May 2026
Safety

PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection

DGX agent

arXiv:2509.26272v3 Announce Type: replace Abstract: The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the

safetyarxiv-cs-cv
11 May 2026
Safety

So many bots but honestly decent alignment

DGX agent

This post likely discusses the proliferation of AI bots and chatbot applications while arguing that despite their abundance, many demonstrate reasonable safety practices or alignment with human values

safetyswyx--x
10 May 2026
Safety

From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation

DGX agent

arXiv:2605.04327v1 Announce Type: new Abstract: We propose an architecture for integrating high-level, human-provided safety rules and operator-aligned semantic preferences into autonomous robot navig

safetyarxiv-cs-ro
7 May 2026
Safety

Hey @bloomberg could you update this? It’s already way out of date. 🙏

DGX agent

Gary Marcus posted on X (formerly Twitter) requesting that Bloomberg update outdated information, likely regarding AI safety, regulation, or related technological topics that Marcus frequently discuss

safetygary-marcus--x
7 May 2026
Safety

LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy

DGX agent

arXiv:2605.04295v1 Announce Type: new Abstract: LLMs' overconfidence, particularly when hallucinating, poses a significant challenge for the deployment of the models in safety-critical settings and ma

safetyarxiv-cs-lg
7 May 2026
Safety

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate

DGX agent

arXiv:2605.03409v1 Announce Type: new Abstract: We present Robust Agent Compensation (RAC), a log-based recovery paradigm (providing a safety net) implemented through an architectural extension that c

safetyarxiv-cs-ai
7 May 2026
Safety

A Three-Stage Offline SDRE-Based Control Framework for Human Motion Reproduction on a Suspended Bipedal Robot

DGX agent

arXiv:2506.04680v2 Announce Type: replace Abstract: During the development of wearable exoskeletons, evaluations involving human subjects pose inherent safety risks. Therefore, systematic testing is o

safetyarxiv-cs-ro
6 May 2026
Safety

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

DGX agent

arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing

safetyarxiv-cs-ai
6 May 2026
Safety

Efficient Temporal Datalog Materialisation for Composite Event Recognition

DGX agent

arXiv:2605.02488v1 Announce Type: new Abstract: Several applications demand the timely detection of critical situations, such as threats to safety and transparency, over high-velocity streams of symbo

safetyarxiv-cs-ai
6 May 2026
Safety

Human-in-the-Loop Uncertainty Analysis in Self-Adaptive Robots Using LLMs

DGX agent

arXiv:2605.02983v1 Announce Type: new Abstract: Self-adaptive robots operate in dynamic, unpredictable environments where unaddressed uncertainties can lead to safety violations and operational failur

safetyarxiv-cs-ro
6 May 2026
Safety

Instance-Level Costs for Nuanced Classifier Evaluation

DGX agent

arXiv:2605.03135v1 Announce Type: new Abstract: Standard classification treats all errors equally, but in content moderation, medical screening, and safety-critical applications, mistakes on clear-cut

safetyarxiv-cs-lg
6 May 2026
Safety

Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings

DGX agent

arXiv:2605.02908v1 Announce Type: new Abstract: Understanding how textual embeddings contribute to memorization in text-to-image diffusion models is crucial for both interpretability and safety. This

safetyarxiv-cs-cv
6 May 2026
Safety

Mira Murati tells the court that she couldn’t trust Sam Altman’s words

DGX agent

Mira Murati, OpenAI's former CTO, has testified under oath that CEO Sam Altman lied to her about the safety standards for a new AI model. In a video deposition shown during the ongoing Musk v. Altman

safetythe-verge-ai
6 May 2026
Safety

Mira Murati’s testimony is gripping – and what it makes absolutely clear is how utterly wrong most of Twitter was about why Sam was fired. –…

DGX agent

Mira Murati’s testimony is gripping – and what it makes absolutely clear is how utterly wrong most of Twitter was about why Sam was fired. – It had nothing per se to do with AI safety - It had nothing

safetygary-marcus--x
6 May 2026
Safety

Model Routing as a Trust Problem: Route Receipts for Adaptive AI Systems

DGX agent

arXiv:2605.01710v1 Announce Type: new Abstract: AI products often route requests through version aliases, service tiers, tool choices, regional endpoints, fallback rules, or safety handling before res

safetyarxiv-cs-ai
6 May 2026
Local Ai

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs

DGX agent

arXiv:2605.02946v1 Announce Type: new Abstract: Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly

local-aiarxiv-cs-lg
6 May 2026
Safety

Steerable Adversarial Scenario Generation through Test-Time Preference Alignment

DGX agent

arXiv:2509.20102v2 Announce Type: cross Abstract: Adversarial scenario generation is a cost-effective approach for safety assessment of autonomous driving systems. However, existing methods are often

safetyarxiv-cs-ro
6 May 2026
Safety

Zero-Shot Signal Temporal Logic Planning with Disjunctive Branch Selection in Dynamic Semantic Maps

DGX agent

arXiv:2605.01222v1 Announce Type: new Abstract: Signal Temporal Logic (STL) offers verifiable task specifications and is crucial for safety-critical control. Yet STL planning remains challenging: exac

safetyarxiv-cs-ai
6 May 2026
Safety

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

DGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

safetyarxiv-cs-cv
5 May 2026
← Previous
1…2930313233…299
Next →