AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

DGX agent

arXiv:2607.24199v1 Announce Type: new Abstract: Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversi

safetyarxiv-cs-cv
28 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

DGX agent

arXiv:2607.13348v1 Announce Type: new Abstract: Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Pr

model-releasesarxiv-cs-ro
16 Jul 2026
Safety

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

DGX agent

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop

safetyarxiv-cs-ai
16 Jul 2026
Safety

Persona Cartography: Charting Language Model Personality Traits in Weight Space

DGX agent

arXiv:2607.07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decompo

safetyarxiv-cs-ai
10 Jul 2026
Safety

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

DGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

safetyarxiv-cs-ai
8 Jul 2026
Safety

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

DGX agent

arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image im

safetyarxiv-cs-cl
2 Jul 2026
Safety

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

DGX agent

arXiv:2504.12557v3 Announce Type: replace-cross Abstract: Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In m

safetyarxiv-cs-ai
1 Jul 2026
Safety

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

DGX agent

arXiv:2603.05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which int

safetyarxiv-cs-ai
30 Jun 2026
Safety

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

DGX agent

arXiv:2606.29115v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoi

safetyarxiv-cs-ai
30 Jun 2026
Safety

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

DGX agent

arXiv:2606.26151v1 Announce Type: cross Abstract: While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventi

safetyarxiv-cs-ai
26 Jun 2026
Safety

Causality-Based Parametric Control Barrier Function for Safe Multi-Vehicle Interaction

DGX agent

arXiv:2606.25134v1 Announce Type: new Abstract: Safe control has been widely studied in various safety-critical applications, for instance, autonomous driving. In order to ensure the autonomous vehicl

safetyarxiv-cs-ro
25 Jun 2026
Safety

From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

DGX agent

arXiv:2502.02205v4 Announce Type: replace Abstract: The application of deep learning for partial differential equation (PDE)-constrained control is gaining increasing attention. However, existing meth

safetyarxiv-cs-lg
25 Jun 2026
Safety

Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection

DGX agent

arXiv:2512.15503v3 Announce Type: replace-cross Abstract: Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle fo

safetyarxiv-cs-ai
24 Jun 2026
Safety

Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking

DGX agent

arXiv:2606.22278v1 Announce Type: cross Abstract: Ensuring safety of learning-enabled robotic manipulation across diverse embodiments and tasks still requires significant manual engineering. Existing

safetyarxiv-cs-lg
23 Jun 2026
Safety

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics

DGX agent

arXiv:2606.22319v1 Announce Type: cross Abstract: Robotic ultrasound scanning in real clinical environments requires both high-level clinical workflow reasoning and low-level closed-loop execution. Ph

safetyarxiv-cs-cv
23 Jun 2026
Safety

Toward Machine Risk Perception: Integrating Trust Calibration and Precursor-Based Risk Estimation for Humanoid

DGX agent

arXiv:2606.20748v1 Announce Type: new Abstract: Humanoid robots are emerging as co-workers in smart manufacturing, yet their dynamic, human-like movements introduce safety risks that differ fundamenta

safetyarxiv-cs-ro
23 Jun 2026
Safety

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

DGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

safetyarxiv-cs-ai
10 Jun 2026
Safety

Emergent alignment and the projectability of ethical personas

DGX agent

arXiv:2606.09475v1 Announce Type: new Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection'

safetyarxiv-cs-ai
9 Jun 2026
Safety

Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human

DGX agent

arXiv:2606.08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approva

safetyarxiv-cs-ai
9 Jun 2026
Safety

Safe, Fluent and Acceptable Motion Generation and Execution for Human--Robot Interaction in Manufacturing Environments

DGX agent

arXiv:2606.08741v1 Announce Type: new Abstract: Robots operating in human environments must not only ensure physical safety but also exhibit behaviors that are understandable, fluent, and acceptable t

safetyarxiv-cs-ro
9 Jun 2026
Safety

Conformal Risk-Averse Decision Making with Action Conditional Guarantee

DGX agent

arXiv:2606.05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety g

safetyarxiv-cs-ai
6 Jun 2026
Safety

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

DGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

safetyarxiv-cs-ai
4 Jun 2026
Safety

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

DGX agent

arXiv:2606.01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may

safetyarxiv-cs-lg
2 Jun 2026
Safety

The Harsh Truth: Segment-Level Analysis of Harsh Driving Events in Milan Using Large-Scale Telematics, Street Networks, and Google Street View

DGX agent

arXiv:2606.00261v1 Announce Type: new Abstract: Police-reported crash statistics remain the standard input for urban road-safety assessment, but their incompleteness and reporting lag limit their usef

safetyarxiv-cs-cv
2 Jun 2026
Safety

Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models

DGX agent

arXiv:2604.17176v2 Announce Type: replace-cross Abstract: Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajecto

safetyarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Safety

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

DGX agent

arXiv:2605.27766v1 Announce Type: new Abstract: LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongsi

safetyarxiv-cs-ai
28 May 2026
Safety

No Safe Dose: How Training Data Drives Unsafe Image Generation

DGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

safetyarxiv-cs-cv
28 May 2026
Safety

Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning

DGX agent

arXiv:2605.26452v1 Announce Type: cross Abstract: Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints duri

safetyarxiv-cs-lg
27 May 2026
Safety

CBANet: A Compact Attention-Based CNN-BiLSTM Network for Aggressive Driving Event Detection

DGX agent

arXiv:2605.23471v1 Announce Type: cross Abstract: Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promisin

safetyarxiv-cs-ai
25 May 2026
Safety

MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems

DGX agent

arXiv:2602.04431v2 Announce Type: replace Abstract: LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fai

safetyarxiv-cs-lg
25 May 2026
Safety

Kernel-Based Safe Exploration in Deep Reinforcement Learning

DGX agent

arXiv:2605.22207v1 Announce Type: cross Abstract: Safety has been a major concern when deploying deep reinforcement learning algorithms in the real world. A promising direction that ensures that the l

safetyarxiv-cs-lg
23 May 2026
Safety

Automatically Learning Construction Injury Precursors from Text

DGX agent

arXiv:1907.11769v4 Announce Type: replace Abstract: In light of the increasing availability of digitally recorded safety reports in the construction industry, it is important to develop methods to exp

safetyarxiv-cs-cl
21 May 2026
Safety

Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions

DGX agent

arXiv:2605.20566v1 Announce Type: new Abstract: Active perception in uncertain environments requires robots to navigate safely while acquiring informative observations to reduce map uncertainty. These

safetyarxiv-cs-ro
21 May 2026
Safety

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

DGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

safetyarxiv-cs-ai
19 May 2026
Safety

Scenario Generation in Roundabouts with Adjustable Interaction Intensity

DGX agent

arXiv:2605.18026v1 Announce Type: new Abstract: Roundabouts, characterized by frequent merging and yielding interactions, remain a safety-critical corner case for the development and testing of intell

safetyarxiv-cs-ro
19 May 2026
Safety

PCASim: Promptable Closed-loop Adversarial Simulation for Urban Traffic Environment

DGX agent

arXiv:2605.15654v1 Announce Type: new Abstract: Real-world autonomous driving, particularly in urban environments with numerous corner cases, requires rigorous testing to ensure product safety and rob

safetyarxiv-cs-ro
18 May 2026
Safety

SLIP & ETHICS: Graduated Intervention for AI Emotional Companions

DGX agent

arXiv:2605.15915v1 Announce Type: cross Abstract: AI emotional companions face a safety-rapport paradox: restrictive safeguards can damage supportive alliance, while permissive systems risk user harm.

safetyarxiv-cs-ai
18 May 2026
Safety

Training ML Models with Predictable Failures

DGX agent

arXiv:2605.15134v1 Announce Type: new Abstract: Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely l

safetyarxiv-cs-lg
15 May 2026
Safety

DisaBench: A Participatory Evaluation Framework for Disability Harms in Language Models

DGX agent

arXiv:2605.12702v1 Announce Type: new Abstract: General-purpose safety benchmarks for large language models do not adequately evaluate disability-related harms. We introduce DisaBench: a taxonomy of t

safetyarxiv-cs-ai
14 May 2026
Safety

Revealing Interpretable Failure Modes of VLMs

DGX agent

arXiv:2605.12674v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to general

safetyarxiv-cs-ai
14 May 2026
Safety

ASACK : Adaptive Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees

DGX agent

arXiv:2605.09659v1 Announce Type: new Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear operator acting on lifted observables, enabli

safetyarxiv-cs-ro
12 May 2026
Safety

Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations

DGX agent

arXiv:2605.08754v1 Announce Type: new Abstract: Taxiway routing and on-surface conflict avoidance are coupled safety-critical decision problems in airport surface operations. Existing planning and opt

safetyarxiv-cs-ai
12 May 2026
Safety

InvThink: Premortem Reasoning for Safer Language Models

DGX agent

arXiv:2510.01569v3 Announce Type: replace Abstract: We present InvThink, a training and prompting framework that requires the model to enumerate, analyze, and constrain potential failures before gener

safetyarxiv-cs-ai
11 May 2026
Safety

Brainrot: Deskilling and Addiction are Overlooked AI Risks

DGX agent

arXiv:2605.03512v1 Announce Type: cross Abstract: The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discri

safetyarxiv-cs-ai
7 May 2026
Safety

Practical validation of synthetic pre-crash scenarios

DGX agent

arXiv:2605.04564v1 Announce Type: new Abstract: The representativeness of synthetic pre-crash scenarios is crucial for assessing the safety impact of Driving Automation Systems through virtual simulat

safetyarxiv-cs-ro
7 May 2026
Safety

Lateral String Stability for Vehicle Platoons: Formulation, Definition, and Analysis

DGX agent

arXiv:2605.01731v1 Announce Type: new Abstract: Platooning of connected and automated vehicles provides significant benefits in terms of energy efficiency, traffic throughput, and, most critically, sa

safetyarxiv-cs-ro
5 May 2026
Safety

Risk Reporting for Developers' Internal AI Model Use

DGX agent

arXiv:2604.24966v1 Announce Type: cross Abstract: Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a p

safetyarxiv-cs-ai
30 Apr 2026
← Previous
1…1617181920…257
Next →