AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

LLM Scheming Inversely Scales with Pretraining Language Coverage

DGX agent

arXiv:2607.24769v1 Announce Type: new Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has emp

safetyarxiv-cs-ai
29 Jul 2026
Safety

Constrained Reinforcement Learning Using Successor Representations

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to int

safetyarxiv-cs-lg
28 Jul 2026
Safety

Physical AI Governance: From Theory to Practice Across Life Cycle

DGX agent

arXiv:2607.22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact wit

safetyarxiv-cs-ai
28 Jul 2026
Safety

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

DGX agent

arXiv:2607.24199v1 Announce Type: new Abstract: Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversi

safetyarxiv-cs-cv
28 Jul 2026
Safety

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and…

DGX agent

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety an

safetyclem-delangue--x
24 Jul 2026
Model Releases

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

DGX agent

arXiv:2607.13348v1 Announce Type: new Abstract: Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Pr

model-releasesarxiv-cs-ro
16 Jul 2026
Safety

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

DGX agent

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop

safetyarxiv-cs-ai
16 Jul 2026
Safety

Persona Cartography: Charting Language Model Personality Traits in Weight Space

DGX agent

arXiv:2607.07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decompo

safetyarxiv-cs-ai
10 Jul 2026
Safety

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

DGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

safetyarxiv-cs-ai
8 Jul 2026
Safety

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

DGX agent

arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image im

safetyarxiv-cs-cl
2 Jul 2026
Safety

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

DGX agent

arXiv:2504.12557v3 Announce Type: replace-cross Abstract: Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In m

safetyarxiv-cs-ai
1 Jul 2026
Safety

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

DGX agent

arXiv:2603.05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which int

safetyarxiv-cs-ai
30 Jun 2026
Safety

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

DGX agent

arXiv:2606.29115v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoi

safetyarxiv-cs-ai
30 Jun 2026
Safety

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

DGX agent

arXiv:2606.26151v1 Announce Type: cross Abstract: While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventi

safetyarxiv-cs-ai
26 Jun 2026
Safety

Causality-Based Parametric Control Barrier Function for Safe Multi-Vehicle Interaction

DGX agent

arXiv:2606.25134v1 Announce Type: new Abstract: Safe control has been widely studied in various safety-critical applications, for instance, autonomous driving. In order to ensure the autonomous vehicl

safetyarxiv-cs-ro
25 Jun 2026
Safety

From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

DGX agent

arXiv:2502.02205v4 Announce Type: replace Abstract: The application of deep learning for partial differential equation (PDE)-constrained control is gaining increasing attention. However, existing meth

safetyarxiv-cs-lg
25 Jun 2026
Safety

Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection

DGX agent

arXiv:2512.15503v3 Announce Type: replace-cross Abstract: Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle fo

safetyarxiv-cs-ai
24 Jun 2026
Safety

Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking

DGX agent

arXiv:2606.22278v1 Announce Type: cross Abstract: Ensuring safety of learning-enabled robotic manipulation across diverse embodiments and tasks still requires significant manual engineering. Existing

safetyarxiv-cs-lg
23 Jun 2026
Safety

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics

DGX agent

arXiv:2606.22319v1 Announce Type: cross Abstract: Robotic ultrasound scanning in real clinical environments requires both high-level clinical workflow reasoning and low-level closed-loop execution. Ph

safetyarxiv-cs-cv
23 Jun 2026
Safety

Toward Machine Risk Perception: Integrating Trust Calibration and Precursor-Based Risk Estimation for Humanoid

DGX agent

arXiv:2606.20748v1 Announce Type: new Abstract: Humanoid robots are emerging as co-workers in smart manufacturing, yet their dynamic, human-like movements introduce safety risks that differ fundamenta

safetyarxiv-cs-ro
23 Jun 2026
Safety

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

DGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

safetyarxiv-cs-ai
10 Jun 2026
Safety

Emergent alignment and the projectability of ethical personas

DGX agent

arXiv:2606.09475v1 Announce Type: new Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection'

safetyarxiv-cs-ai
9 Jun 2026
Safety

Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human

DGX agent

arXiv:2606.08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approva

safetyarxiv-cs-ai
9 Jun 2026
Safety

Safe, Fluent and Acceptable Motion Generation and Execution for Human--Robot Interaction in Manufacturing Environments

DGX agent

arXiv:2606.08741v1 Announce Type: new Abstract: Robots operating in human environments must not only ensure physical safety but also exhibit behaviors that are understandable, fluent, and acceptable t

safetyarxiv-cs-ro
9 Jun 2026
Safety

Conformal Risk-Averse Decision Making with Action Conditional Guarantee

DGX agent

arXiv:2606.05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety g

safetyarxiv-cs-ai
6 Jun 2026
Safety

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

DGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

safetyarxiv-cs-ai
4 Jun 2026
Safety

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

DGX agent

arXiv:2606.01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may

safetyarxiv-cs-lg
2 Jun 2026
Safety

The Harsh Truth: Segment-Level Analysis of Harsh Driving Events in Milan Using Large-Scale Telematics, Street Networks, and Google Street View

DGX agent

arXiv:2606.00261v1 Announce Type: new Abstract: Police-reported crash statistics remain the standard input for urban road-safety assessment, but their incompleteness and reporting lag limit their usef

safetyarxiv-cs-cv
2 Jun 2026
Safety

Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models

DGX agent

arXiv:2604.17176v2 Announce Type: replace-cross Abstract: Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajecto

safetyarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Safety

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

DGX agent

arXiv:2605.27766v1 Announce Type: new Abstract: LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongsi

safetyarxiv-cs-ai
28 May 2026
Safety

No Safe Dose: How Training Data Drives Unsafe Image Generation

DGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

safetyarxiv-cs-cv
28 May 2026
Safety

Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning

DGX agent

arXiv:2605.26452v1 Announce Type: cross Abstract: Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints duri

safetyarxiv-cs-lg
27 May 2026
Safety

CBANet: A Compact Attention-Based CNN-BiLSTM Network for Aggressive Driving Event Detection

DGX agent

arXiv:2605.23471v1 Announce Type: cross Abstract: Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promisin

safetyarxiv-cs-ai
25 May 2026
Safety

MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems

DGX agent

arXiv:2602.04431v2 Announce Type: replace Abstract: LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fai

safetyarxiv-cs-lg
25 May 2026
Safety

Kernel-Based Safe Exploration in Deep Reinforcement Learning

DGX agent

arXiv:2605.22207v1 Announce Type: cross Abstract: Safety has been a major concern when deploying deep reinforcement learning algorithms in the real world. A promising direction that ensures that the l

safetyarxiv-cs-lg
23 May 2026
Safety

Automatically Learning Construction Injury Precursors from Text

DGX agent

arXiv:1907.11769v4 Announce Type: replace Abstract: In light of the increasing availability of digitally recorded safety reports in the construction industry, it is important to develop methods to exp

safetyarxiv-cs-cl
21 May 2026
Safety

Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions

DGX agent

arXiv:2605.20566v1 Announce Type: new Abstract: Active perception in uncertain environments requires robots to navigate safely while acquiring informative observations to reduce map uncertainty. These

safetyarxiv-cs-ro
21 May 2026
Safety

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

DGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

safetyarxiv-cs-ai
19 May 2026
Safety

Scenario Generation in Roundabouts with Adjustable Interaction Intensity

DGX agent

arXiv:2605.18026v1 Announce Type: new Abstract: Roundabouts, characterized by frequent merging and yielding interactions, remain a safety-critical corner case for the development and testing of intell

safetyarxiv-cs-ro
19 May 2026
Safety

PCASim: Promptable Closed-loop Adversarial Simulation for Urban Traffic Environment

DGX agent

arXiv:2605.15654v1 Announce Type: new Abstract: Real-world autonomous driving, particularly in urban environments with numerous corner cases, requires rigorous testing to ensure product safety and rob

safetyarxiv-cs-ro
18 May 2026
Safety

SLIP & ETHICS: Graduated Intervention for AI Emotional Companions

DGX agent

arXiv:2605.15915v1 Announce Type: cross Abstract: AI emotional companions face a safety-rapport paradox: restrictive safeguards can damage supportive alliance, while permissive systems risk user harm.

safetyarxiv-cs-ai
18 May 2026
Safety

Training ML Models with Predictable Failures

DGX agent

arXiv:2605.15134v1 Announce Type: new Abstract: Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely l

safetyarxiv-cs-lg
15 May 2026
Safety

DisaBench: A Participatory Evaluation Framework for Disability Harms in Language Models

DGX agent

arXiv:2605.12702v1 Announce Type: new Abstract: General-purpose safety benchmarks for large language models do not adequately evaluate disability-related harms. We introduce DisaBench: a taxonomy of t

safetyarxiv-cs-ai
14 May 2026
Safety

Revealing Interpretable Failure Modes of VLMs

DGX agent

arXiv:2605.12674v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to general

safetyarxiv-cs-ai
14 May 2026
Safety

Smart moves: Building resilient transportation systems with Google AI

DGX agent

What does transportation mean to you? For some, it’s making sure the train is on schedule so they can get to work on time. Maybe it’s making sure you have time connecting between flights. Maybe it’s a

safetygoogle-cloud-ai
13 May 2026
Safety

ASACK : Adaptive Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees

DGX agent

arXiv:2605.09659v1 Announce Type: new Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear operator acting on lifted observables, enabli

safetyarxiv-cs-ro
12 May 2026
Safety

Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations

DGX agent

arXiv:2605.08754v1 Announce Type: new Abstract: Taxiway routing and on-surface conflict avoidance are coupled safety-critical decision problems in airport surface operations. Existing planning and opt

safetyarxiv-cs-ai
12 May 2026
← Previous
1…1819202122…297
Next →