AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Safety

Graph-Guided Safe Diffuser: Topological Graph Guidance for Safe Diffusion Planning

DGX agent

arXiv:2608.09484v1 Announce Type: new Abstract: Many diffusion-based planners enforce safety through inference-time guidance, but such interleaved trajectory deformations often degrade kinematic feasi

safetyarxiv-cs-ro
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

DGX agent

arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Howev

safetyarxiv-cs-ai
11 Aug 2026
Safety

Forbidden Region Dynamic Active Constraints in Robot-Assisted Minimally Invasive Surgery

DGX agent

arXiv:2608.03010v1 Announce Type: new Abstract: In robot-assisted surgery, Forbidden Region Active Constraints (FRAC) represent a control strategy that helps maintain task safety by generating anisotr

safetyarxiv-cs-ro
5 Aug 2026
Safety

Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling

DGX agent

arXiv:2607.27511v1 Announce Type: new Abstract: Imitation learning has shown increasing promise for autonomous robotic surgery, yet safe deployment remains challenging due to the safety-critical natur

safetyarxiv-cs-ro
31 Jul 2026
Safety

Uncertainty quantification for trustworthy deep learning: Methods and measures

DGX agent

arXiv:2607.28248v1 Announce Type: cross Abstract: The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures

safetyarxiv-cs-lg
31 Jul 2026
Safety

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

DGX agent

arXiv:2607.26574v1 Announce Type: cross Abstract: Safety classifiers ('guards') are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning:

safetyarxiv-cs-lg
30 Jul 2026
Safety

Steering Instruction Hierarchies at Inference Time

DGX agent

arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override confl

safetyarxiv-cs-cl
30 Jul 2026
Safety

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, man…

DGX agent

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, many forefront members of the AI-Safety community, in their fer

safetyyann-lecun--x
30 Jul 2026
Safety

Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail

DGX agent

arXiv:2607.25664v1 Announce Type: new Abstract: Machine learning demand forecasts optimize statistical accuracy yet leave excess operational volatility that inflates safety stock and amplifies the Bul

safetyarxiv-cs-lg
29 Jul 2026
Safety

LLM Scheming Inversely Scales with Pretraining Language Coverage

DGX agent

arXiv:2607.24769v1 Announce Type: new Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has emp

safetyarxiv-cs-ai
29 Jul 2026
Safety

Constrained Reinforcement Learning Using Successor Representations

DGX agent

arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to int

safetyarxiv-cs-lg
28 Jul 2026
Safety

Physical AI Governance: From Theory to Practice Across Life Cycle

DGX agent

arXiv:2607.22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact wit

safetyarxiv-cs-ai
28 Jul 2026
Safety

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

DGX agent

arXiv:2607.24199v1 Announce Type: new Abstract: Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversi

safetyarxiv-cs-cv
28 Jul 2026
Safety

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and…

DGX agent

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety an

safetyclem-delangue--x
24 Jul 2026
Model Releases

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

DGX agent

arXiv:2607.13348v1 Announce Type: new Abstract: Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Pr

model-releasesarxiv-cs-ro
16 Jul 2026
Safety

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

DGX agent

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop

safetyarxiv-cs-ai
16 Jul 2026
Safety

Persona Cartography: Charting Language Model Personality Traits in Weight Space

DGX agent

arXiv:2607.07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decompo

safetyarxiv-cs-ai
10 Jul 2026
Safety

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

DGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

safetyarxiv-cs-ai
8 Jul 2026
Safety

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

DGX agent

arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image im

safetyarxiv-cs-cl
2 Jul 2026
Safety

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

DGX agent

arXiv:2504.12557v3 Announce Type: replace-cross Abstract: Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In m

safetyarxiv-cs-ai
1 Jul 2026
Safety

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

DGX agent

arXiv:2603.05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which int

safetyarxiv-cs-ai
30 Jun 2026
Safety

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

DGX agent

arXiv:2606.29115v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoi

safetyarxiv-cs-ai
30 Jun 2026
Safety

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

DGX agent

arXiv:2606.26151v1 Announce Type: cross Abstract: While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventi

safetyarxiv-cs-ai
26 Jun 2026
Safety

Causality-Based Parametric Control Barrier Function for Safe Multi-Vehicle Interaction

DGX agent

arXiv:2606.25134v1 Announce Type: new Abstract: Safe control has been widely studied in various safety-critical applications, for instance, autonomous driving. In order to ensure the autonomous vehicl

safetyarxiv-cs-ro
25 Jun 2026
Safety

From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

DGX agent

arXiv:2502.02205v4 Announce Type: replace Abstract: The application of deep learning for partial differential equation (PDE)-constrained control is gaining increasing attention. However, existing meth

safetyarxiv-cs-lg
25 Jun 2026
Safety

Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection

DGX agent

arXiv:2512.15503v3 Announce Type: replace-cross Abstract: Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle fo

safetyarxiv-cs-ai
24 Jun 2026
Safety

Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking

DGX agent

arXiv:2606.22278v1 Announce Type: cross Abstract: Ensuring safety of learning-enabled robotic manipulation across diverse embodiments and tasks still requires significant manual engineering. Existing

safetyarxiv-cs-lg
23 Jun 2026
Safety

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics

DGX agent

arXiv:2606.22319v1 Announce Type: cross Abstract: Robotic ultrasound scanning in real clinical environments requires both high-level clinical workflow reasoning and low-level closed-loop execution. Ph

safetyarxiv-cs-cv
23 Jun 2026
Safety

Toward Machine Risk Perception: Integrating Trust Calibration and Precursor-Based Risk Estimation for Humanoid

DGX agent

arXiv:2606.20748v1 Announce Type: new Abstract: Humanoid robots are emerging as co-workers in smart manufacturing, yet their dynamic, human-like movements introduce safety risks that differ fundamenta

safetyarxiv-cs-ro
23 Jun 2026
Safety

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

DGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

safetyarxiv-cs-ai
10 Jun 2026
Safety

Emergent alignment and the projectability of ethical personas

DGX agent

arXiv:2606.09475v1 Announce Type: new Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection'

safetyarxiv-cs-ai
9 Jun 2026
Safety

Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human

DGX agent

arXiv:2606.08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approva

safetyarxiv-cs-ai
9 Jun 2026
Safety

Safe, Fluent and Acceptable Motion Generation and Execution for Human--Robot Interaction in Manufacturing Environments

DGX agent

arXiv:2606.08741v1 Announce Type: new Abstract: Robots operating in human environments must not only ensure physical safety but also exhibit behaviors that are understandable, fluent, and acceptable t

safetyarxiv-cs-ro
9 Jun 2026
Safety

Conformal Risk-Averse Decision Making with Action Conditional Guarantee

DGX agent

arXiv:2606.05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety g

safetyarxiv-cs-ai
6 Jun 2026
Safety

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

DGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

safetyarxiv-cs-ai
4 Jun 2026
Safety

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

DGX agent

arXiv:2606.01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may

safetyarxiv-cs-lg
2 Jun 2026
Safety

The Harsh Truth: Segment-Level Analysis of Harsh Driving Events in Milan Using Large-Scale Telematics, Street Networks, and Google Street View

DGX agent

arXiv:2606.00261v1 Announce Type: new Abstract: Police-reported crash statistics remain the standard input for urban road-safety assessment, but their incompleteness and reporting lag limit their usef

safetyarxiv-cs-cv
2 Jun 2026
Safety

Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models

DGX agent

arXiv:2604.17176v2 Announce Type: replace-cross Abstract: Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajecto

safetyarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Safety

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

DGX agent

arXiv:2605.27766v1 Announce Type: new Abstract: LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongsi

safetyarxiv-cs-ai
28 May 2026
Safety

No Safe Dose: How Training Data Drives Unsafe Image Generation

DGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

safetyarxiv-cs-cv
28 May 2026
Safety

Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning

DGX agent

arXiv:2605.26452v1 Announce Type: cross Abstract: Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints duri

safetyarxiv-cs-lg
27 May 2026
Safety

CBANet: A Compact Attention-Based CNN-BiLSTM Network for Aggressive Driving Event Detection

DGX agent

arXiv:2605.23471v1 Announce Type: cross Abstract: Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promisin

safetyarxiv-cs-ai
25 May 2026
Safety

MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems

DGX agent

arXiv:2602.04431v2 Announce Type: replace Abstract: LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fai

safetyarxiv-cs-lg
25 May 2026
Safety

Kernel-Based Safe Exploration in Deep Reinforcement Learning

DGX agent

arXiv:2605.22207v1 Announce Type: cross Abstract: Safety has been a major concern when deploying deep reinforcement learning algorithms in the real world. A promising direction that ensures that the l

safetyarxiv-cs-lg
23 May 2026
Safety

Automatically Learning Construction Injury Precursors from Text

DGX agent

arXiv:1907.11769v4 Announce Type: replace Abstract: In light of the increasing availability of digitally recorded safety reports in the construction industry, it is important to develop methods to exp

safetyarxiv-cs-cl
21 May 2026
Safety

Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions

DGX agent

arXiv:2605.20566v1 Announce Type: new Abstract: Active perception in uncertain environments requires robots to navigate safely while acquiring informative observations to reduce map uncertainty. These

safetyarxiv-cs-ro
21 May 2026
Safety

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

DGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

safetyarxiv-cs-ai
19 May 2026
← Previous
1…1819202122…299
Next →