AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

CAREBench: A Child-Safety Risk Benchmark for Language Models

DGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

model-releasesarxiv-cs-lg
30 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Algorithms for Deciding the Safety of States in Fully Observable Non-deterministic Problems: Technical Report

DGX agent

arXiv:2603.15282v2 Announce Type: replace Abstract: Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees. Recent work introduced

safetyarxiv-cs-ai
29 Jun 2026
Safety

Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations

DGX agent

arXiv:2606.25629v1 Announce Type: new Abstract: Robot navigation in safety-critical scenarios faces significant challenges from unforeseen semantic events, where collisions arise primarily from the un

safetyarxiv-cs-ro
25 Jun 2026
Safety

Verifiable Foundation Models for Robot Safety

DGX agent

arXiv:2606.23754v1 Announce Type: cross Abstract: Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes the

safetyarxiv-cs-lg
24 Jun 2026
Safety

How Well Do Latent World Models Understand Partially Observable Safety Constraints?

DGX agent

arXiv:2510.06492v2 Announce Type: replace Abstract: Latent world models are a promising approach for learning state representations and dynamics directly from high-dimensional observations, enabling r

safetyarxiv-cs-ro
9 Jun 2026
Safety

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

DGX agent

arXiv:2606.09165v1 Announce Type: new Abstract: Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle u

safetyarxiv-cs-ai
9 Jun 2026
Safety

Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models

DGX agent

arXiv:2606.08451v1 Announce Type: cross Abstract: Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy. Althoug

safetyarxiv-cs-ai
9 Jun 2026
Safety

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

DGX agent

arXiv:2606.08531v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, a

safetyarxiv-cs-ai
9 Jun 2026
Safety

Learning of Robot Safety Policies via Adversarial Synthetic Scenarios

DGX agent

arXiv:2606.05952v1 Announce Type: new Abstract: In this work, we propose an agentic gamification framework for hazard-informed learning of robot safety policies through synthetic scenarios. We model s

safetyarxiv-cs-ro
5 Jun 2026
Safety

RiskFlow: Fast and Faithful Safety-Critical Traffic Scenario Generation

DGX agent

arXiv:2606.06423v1 Announce Type: new Abstract: Safety-critical traffic scenario generation is essential for evaluating autonomous driving systems under rare but high-risk interactions. Existing diffu

safetyarxiv-cs-ro
5 Jun 2026
Model Releases

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

DGX agent

arXiv:2606.03954v1 Announce Type: new Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible conse

model-releasesarxiv-cs-cv
3 Jun 2026
Safety

MESA: Improving MoE Safety Alignment via Decentralized Expertise

DGX agent

arXiv:2606.00651v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures scale Large Language Models (LLMs) efficiently, enabling greater capacity with reduced computational cost by dy

safetyarxiv-cs-ai
2 Jun 2026
Safety

Safety Alignment of LMs via Non-cooperative Games

DGX agent

arXiv:2512.20806v3 Announce Type: replace Abstract: Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rel

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

DGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

DGX agent

arXiv:2605.28830v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a c

model-releasesarxiv-cs-ai
29 May 2026
Safety

SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving

DGX agent

arXiv:2605.28583v1 Announce Type: cross Abstract: Ensuring both safety and efficiency in decision-making for autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement

safetyarxiv-cs-ai
28 May 2026
Safety

Curriculum Learning for Safety Alignment

DGX agent

arXiv:2605.26315v1 Announce Type: cross Abstract: Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibit

safetyarxiv-cs-ai
27 May 2026
Safety

AI-based Prediction of Independent Construction Safety Outcomes from Universal Attributes

DGX agent

arXiv:1908.05972v3 Announce Type: replace Abstract: This paper significantly improves on, and finishes to validate, an approach proposed in previous research in which safety outcomes were predicted fr

safetyarxiv-cs-lg
21 May 2026
Safety

PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment

DGX agent

arXiv:2605.21225v1 Announce Type: new Abstract: We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it f

safetyarxiv-cs-lg
21 May 2026
Safety

Distributional AGI Safety

DGX agent

arXiv:2512.16856v2 Announce Type: replace Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an e

safetyarxiv-cs-ai
20 May 2026
Safety

Reactive Robot-Centric Safety for Autonomous Navigation in Constrained and Dynamic Environments

DGX agent

arXiv:2605.15782v1 Announce Type: new Abstract: In this work, we address the problem of ensuring real-time safety in autonomous robot navigation, in spatially constrained dynamic environments, by util

safetyarxiv-cs-ro
18 May 2026
Model Releases

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

DGX agent

arXiv:2605.12386v1 Announce Type: new Abstract: Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are tempo

model-releasesarxiv-cs-ro
13 May 2026
Safety

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

DGX agent

arXiv:2601.18061v3 Announce Type: replace Abstract: Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI

safetyarxiv-cs-ai
12 May 2026
Safety

GuardAD: Safeguarding Autonomous Driving MLLMs via Markovian Safety Logic

DGX agent

arXiv:2605.10386v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse sa

safetyarxiv-cs-ai
12 May 2026
Safety

Why Does Agentic Safety Fail to Generalize Across Tasks?

DGX agent

arXiv:2605.06992v1 Announce Type: new Abstract: AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen

safetyarxiv-cs-lg
11 May 2026
Safety

Safety by Invariance, Liveness through Refinement: Heterogeneous Contract Framework for Co-Design of Layered Control

DGX agent

arXiv:2605.04222v1 Announce Type: cross Abstract: Real-world control systems must achieve long-horizon objectives (liveness) while respecting continuous-time safety constraints, a combination that mot

safetyarxiv-cs-ro
7 May 2026
Safety

Safety-critical Control Under Partial Observability: Reach-Avoid POMDP meets Belief Space Control

DGX agent

arXiv:2603.10572v2 Announce Type: replace Abstract: Partially Observable Markov Decision Processes (POMDPs) provide a principled framework for robot decision-making under uncertainty. Solving reach-av

safetyarxiv-cs-ro
6 May 2026
Safety

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

DGX agent

arXiv:2605.00227v1 Announce Type: new Abstract: There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely

safetyarxiv-cs-cl
4 May 2026
Safety

Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification

DGX agent

arXiv:2605.00326v1 Announce Type: new Abstract: Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they a

safetyarxiv-cs-cl
4 May 2026
Safety

Focus Session: Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification

DGX agent

arXiv:2604.27807v1 Announce Type: new Abstract: The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by

safetyarxiv-cs-ai
1 May 2026
Safety

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

DGX agent

arXiv:2604.28087v1 Announce Type: cross Abstract: Rule-based systems remain central in safety-critical domains but often struggle with scalability, brittleness, and goal misspecification. These limita

safetyarxiv-cs-ai
1 May 2026
Safety

A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

DGX agent

arXiv:2505.10924v4 Announce Type: replace-cross Abstract: Recently, AI-driven interactions with computing devices have advanced from basic prototype tools to sophisticated, LLM-based systems that emul

safetyarxiv-cs-ai
30 Apr 2026
Safety

Recipes for Calibration Checks in Safety-Critical Applications

DGX agent

arXiv:2604.26479v1 Announce Type: cross Abstract: Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters

safetyarxiv-cs-lg
30 Apr 2026
Safety

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

DGX agent

arXiv:2509.26238v4 Announce Type: replace Abstract: Monitoring large language models' (LLMs) activations is an effective way to detect harmful requests before they lead to unsafe outputs. However, tra

safetyarxiv-cs-lg
22 Apr 2026
Model Releases

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

DGX agent

arXiv:2505.15404v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enha

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

DGX agent

arXiv:2604.17691v1 Announce Type: new Abstract: Safety alignment in large language models is remarkably shallow: it is concentrated in the first few output tokens and reversible by fine-tuning on as f

model-releasesarxiv-cs-lg
21 Apr 2026
Local Ai

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

DGX agent

arXiv:2604.16542v1 Announce Type: cross Abstract: Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). How

local-aiarxiv-cs-cl
21 Apr 2026
Safety

Trajectory Planning for Safe Dual Control with Active Exploration

DGX agent

arXiv:2604.15507v1 Announce Type: new Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations, ye

safetyarxiv-cs-ro
20 Apr 2026
Safety

Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades

DGX agent

arXiv:2604.14251v1 Announce Type: new Abstract: Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalat

safetyarxiv-cs-lg
17 Apr 2026
Safety

ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance

DGX agent

arXiv:2604.12308v1 Announce Type: new Abstract: Individuals' concerns about data privacy and AI safety are highly contextualized and extend beyond sensitive patterns. Addressing these issues requires

safetyarxiv-cs-cl
15 Apr 2026
Model Releases

LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

DGX agent

arXiv:2604.12710v1 Announce Type: cross Abstract: Large language models (LLMs) often demonstrate strong safety performance in high-resource languages, yet exhibit severe vulnerabilities when queried i

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

LLM-based Realistic Safety-Critical Driving Video Generation

DGX agent

arXiv:2507.01264v2 Announce Type: replace-cross Abstract: Designing diverse and safety-critical driving scenarios is essential for evaluating autonomous driving systems. In this paper, we propose a no

safetyarxiv-cs-ai
14 Apr 2026
Safety

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

DGX agent

arXiv:2602.03402v3 Announce Type: replace Abstract: Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerabl

safetyarxiv-cs-ai
14 Apr 2026
Safety

Safety Guarantees in Zero-Shot Reinforcement Learning for Cascade Dynamical Systems

DGX agent

arXiv:2604.10429v1 Announce Type: new Abstract: This paper considers the problem of zero-shot safety guarantees for cascade dynamical systems. These are systems where a subset of the states (the inner

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

DGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

Safe and Nonconservative Contingency Planning for Autonomous Vehicles via Online Learning-Based Reachable Set Barriers

DGX agent

arXiv:2509.07464v2 Announce Type: replace Abstract: Autonomous vehicles must navigate dynamically uncertain environments while balancing safety and efficiency. This challenge is exacerbated by unpredi

safetyarxiv-cs-ro
16 Apr 2026
Safety

The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software

DGX agent

arXiv:2608.10025v1 Announce Type: new Abstract: For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure events experienced by the software) can

safetyarxiv-cs-ro
12 Aug 2026
Model Releases

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

DGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

model-releasesarxiv-cs-cl
11 Aug 2026
← Previous
1…45678…255
Next →