AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Cohere has joined @NVIDIA alongside industry leaders in founding the Open Secure AI Alliance. Everybody should have the capability to keep t…

DGX agent

Cohere has joined @NVIDIA alongside industry leaders in founding the Open Secure AI Alliance. Everybody should have the capability to keep their infrastructure secure. Everybody deserves access to mod

safetycohere--x
30 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

DGX agent

arXiv:2607.26034v1 Announce Type: new Abstract: Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful.

safetyarxiv-cs-ai
29 Jul 2026
Safety

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

DGX agent

arXiv:2607.25648v1 Announce Type: cross Abstract: Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressur

safetyarxiv-cs-ai
29 Jul 2026
Safety

An Unofficial FastLAS Tutorial: A Programmer's Guide

DGX agent

arXiv:2607.23557v1 Announce Type: cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and

safetyarxiv-cs-ai
28 Jul 2026
Safety

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

DGX agent

arXiv:2607.22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. H

safetyarxiv-cs-ai
28 Jul 2026
Safety

Unifying Complementarity Constraints and Control Barrier Functions for Safe Whole-Body Robot Control

DGX agent

arXiv:2504.17647v2 Announce Type: replace Abstract: Safety-critical whole-body robot control demands reactive methods that ensure collision avoidance in real-time. Complementarity constraints and cont

safetyarxiv-cs-ro
28 Jul 2026
Safety

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models

DGX agent

arXiv:2607.20436v1 Announce Type: cross Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A

safetyarxiv-cs-ai
24 Jul 2026
Safety

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

DGX agent

arXiv:2607.13172v1 Announce Type: new Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown a

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model

DGX agent

arXiv:2603.08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots. Classical navigation approaches offer safety a

model-releasesarxiv-cs-lg
15 Jul 2026
Safety

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

DGX agent

arXiv:2607.12784v1 Announce Type: cross Abstract: Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However

safetyarxiv-cs-lg
15 Jul 2026
Safety

Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering

DGX agent

arXiv:2603.28583v2 Announce Type: replace-cross Abstract: Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structure

safetyarxiv-cs-ai
15 Jul 2026
Safety

INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis

DGX agent

arXiv:2607.08316v1 Announce Type: new Abstract: Vehicle intention prediction is a pivotal aspect in the agility and safety of autonomous vehicles in all driving scenarios; if genuine enhancement of au

safetyarxiv-cs-ai
10 Jul 2026
Safety

LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection

DGX agent

arXiv:2607.06592v1 Announce Type: cross Abstract: Object detectors have many applications in safety-critical systems, but they are known to be sensitive to worst-case perturbations such as adversarial

safetyarxiv-cs-ai
9 Jul 2026
Safety

Explainable Reinforcement Learning for Adaptive Traffic Signal Control

DGX agent

arXiv:2607.03703v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control. However, in safety-critical infrastructure like traf

safetyarxiv-cs-ai
7 Jul 2026
Safety

Hope for the Best, Prepare for the Worst: Occlusion-Aware Contingency Planning for Autonomous Vehicles

DGX agent

arXiv:2607.03155v1 Announce Type: new Abstract: The deployment of autonomous vehicles in urban environments introduces significant safety challenges, particularly in scenarios with occlusions, where c

safetyarxiv-cs-ro
7 Jul 2026
Safety

Resolving Primitive-Sharing Ambiguity in Long-Tailed TLS-Based Industrial MEP Point Cloud Segmentation via Spatial Context Constraints

DGX agent

arXiv:2601.19128v2 Announce Type: replace Abstract: In terrestrial laser scanning (TLS)-based mechanical, electrical, and plumbing (MEP) point cloud segmentation, safety-critical components such as re

safetyarxiv-cs-cv
7 Jul 2026
Safety

F1 in Britain: Automated software to blame for crushing expectations

DGX agent

Race control displayed a 'Safety Car in this lap' message during the 2026 British Grand Prix at Silverstone, raising expectations for a final lap restart that never materialized, as the message was er

safetyars-technica
6 Jul 2026
Safety

Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring

DGX agent

arXiv:2607.02121v1 Announce Type: cross Abstract: As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical. Gu

safetyarxiv-cs-ai
3 Jul 2026
Safety

Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

DGX agent

arXiv:2607.01794v1 Announce Type: cross Abstract: With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection,

safetyarxiv-cs-ai
3 Jul 2026
Safety

Robust Operational Space Control with Conformal Disturbance Bounds for Safe Redundant Manipulation

DGX agent

arXiv:2607.00424v1 Announce Type: new Abstract: Redundant robotic manipulators operating in constrained and human-interactive environments require accurate task-space tracking together with rigorous s

safetyarxiv-cs-ro
2 Jul 2026
Safety

On Optimizing Multimodal Jailbreaks for Spoken Language Models

DGX agent

arXiv:2603.19127v2 Announce Type: replace Abstract: As Spoken Language Models (SLMs) integrate speech and text modalities, they inherit the safety vulnerabilities of their LLM backbone while introduci

safetyarxiv-cs-lg
1 Jul 2026
Safety

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

DGX agent

arXiv:2602.03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrop

safetyarxiv-cs-ai
1 Jul 2026
Safety

A Gravitational Interpretation of Fine-Tuning Reversion

DGX agent

arXiv:2606.28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearne

safetyarxiv-cs-ai
30 Jun 2026
Safety

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

DGX agent

arXiv:2602.21160v4 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a

safetyarxiv-cs-lg
30 Jun 2026
Safety

Sparse Autoencoders are Capable LLM Jailbreak Mitigators

DGX agent

arXiv:2602.12418v2 Announce Type: replace-cross Abstract: Jailbreak attacks remain a persistent threat to large language model safety. We propose Context-Conditioned Delta Steering (CC-Delta), an SAE-

safetyarxiv-cs-cl
30 Jun 2026
Safety

The Undecidability of Artificial General Intelligence (AGI) Alignment

DGX agent

arXiv:2606.28639v1 Announce Type: cross Abstract: This article establishes the foundational mathematical limits of Artificial General Intelligence (AGI) safety, proving that the core barrier is not th

safetyarxiv-cs-ai
30 Jun 2026
Safety

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

DGX agent

arXiv:2606.26205v1 Announce Type: new Abstract: Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records,

safetyarxiv-cs-ai
26 Jun 2026
Safety

A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

DGX agent

arXiv:2606.25380v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural context

safetyarxiv-cs-cl
25 Jun 2026
Safety

ARTOO-DARTU: Studying AR-HRC With AR Obstruction Mitigation During a Warehouse Task

DGX agent

arXiv:2606.25202v1 Announce Type: cross Abstract: Human-robot collaboration (HRC) often requires robot intentions and internal states to be conveyed to users for task efficiency and safety. Recently,

safetyarxiv-cs-ro
25 Jun 2026
Safety

Conformal Recovery-Deadline Certificates for Runtime Assurance of Adapting Controllers

DGX agent

arXiv:2606.25371v1 Announce Type: cross Abstract: Runtime assurance (RTA) protects a safety-critical system by switching from an advanced controller to a verified safe controller when a monitored cond

safetyarxiv-cs-ai
25 Jun 2026
Safety

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

DGX agent

arXiv:2606.20681v1 Announce Type: new Abstract: Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Co

safetyarxiv-cs-cv
23 Jun 2026
Safety

HumanHalo -- Safe and Efficient 3D Navigation Among Humans via Minimally Conservative MPC

DGX agent

arXiv:2510.17525v3 Announce Type: replace Abstract: Safe and efficient robotic navigation among humans is essential for integrating robots into everyday environments. Most existing approaches focus on

safetyarxiv-cs-ro
23 Jun 2026
Model Releases

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

DGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

Can Data Work be Reparative?

DGX agent

arXiv:2606.09408v1 Announce Type: cross Abstract: We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and b

safetyarxiv-cs-ai
9 Jun 2026
Safety

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

DGX agent

arXiv:2606.09178v1 Announce Type: cross Abstract: Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target l

safetyarxiv-cs-ai
9 Jun 2026
Safety

Diffuse AI Control on Fuzzy Tasks

DGX agent

arXiv:2606.08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfiel

safetyarxiv-cs-lg
9 Jun 2026
Safety

A Causal Probabilistic Framework for Perception-Informed Closed-Loop Simulation of Autonomous Driving

DGX agent

arXiv:2606.07186v1 Announce Type: new Abstract: Software-in-the-loop (SIL) simulation is a cornerstone for the validation of modern automotive safety functions. However, many current frameworks utiliz

safetyarxiv-cs-ro
8 Jun 2026
Safety

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

DGX agent

arXiv:2606.07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded

safetyarxiv-cs-ai
8 Jun 2026
Safety

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

DGX agent

arXiv:2606.06875v1 Announce Type: new Abstract: Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the

safetyarxiv-cs-cv
8 Jun 2026
Safety

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

DGX agent

arXiv:2606.05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow

safetyarxiv-cs-ai
6 Jun 2026
Safety

Where Should Knowledge Enter? A Layered Framework for Knowledge Infusion in Multimodal Iterative Generative Mo

DGX agent

arXiv:2606.06356v1 Announce Type: new Abstract: Multimodal generative models produce fluent outputs but remain unreliable when generation must respect structured, domain-specific, or safety-critical k

safetyarxiv-cs-ai
6 Jun 2026
Safety

Drishti AI-Event Guardian: An Intelligent Real-Time Crowd Monitoring and Emergency Response System for Mass Gathering Events

DGX agent

arXiv:2606.05185v1 Announce Type: cross Abstract: Mass gathering events are associated with critical safety incidents caused by insufficient crowd monitoring and inadequate emergency response coordina

safetyarxiv-cs-cv
5 Jun 2026
Safety

Instance-Level Post Hoc Uncertainty Quantification in Object Detection

DGX agent

arXiv:2606.04656v1 Announce Type: cross Abstract: Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for saf

safetyarxiv-cs-ai
4 Jun 2026
Safety

Multi-Agent Next-Best-View Optimization for Risk-Averse Planning

DGX agent

arXiv:2606.04158v1 Announce Type: new Abstract: Multi-agent Next-Best-View (NBV) selection for safe path planning in uncertain and unknown environments requires informative, safety-aware, and efficien

safetyarxiv-cs-ro
4 Jun 2026
Safety

Easy-to-Use Shielding for Reinforcement Learning

DGX agent

arXiv:2606.03804v1 Announce Type: new Abstract: Safe exploration is a key challenge in Reinforcement Learning (RL) that aims to prevent agents from making harmful decisions while exploring their envir

safetyarxiv-cs-lg
3 Jun 2026
Safety

Jailbreak Attack Initializations as Extractors of Compliance Directions

DGX agent

arXiv:2502.09755v4 Announce Type: replace-cross Abstract: Safety-aligned LLMs respond to prompts with either compliance or refusal, each corresponding to distinct directions in the model's activation

safetyarxiv-cs-lg
3 Jun 2026
Safety

Embedding Semantic Risk into Distance Fields and CBFs for Online Monocular Safe Control

DGX agent

arXiv:2606.01605v1 Announce Type: new Abstract: We propose an online monocular perception-to-control framework that embeds semantic risk into the distance field used by Control Barrier Function (CBF)-

safetyarxiv-cs-ro
2 Jun 2026
Safety

Multi-Objective Reinforcement Learning for Tactical Decision Making for Trucks in Highway Traffic

DGX agent

arXiv:2601.18783v2 Announce Type: replace-cross Abstract: Balancing safety, efficiency, and operational costs in highway driving poses a challenging decision-making problem for heavy-duty vehicles. A

safetyarxiv-cs-ai
2 Jun 2026
← Previous
1…2021222324…297
Next →