AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Beyond Safety Filtering: Control Barrier Function-Informed Reinforcement Learning for Connected and Automated Vehicles

DGX agent

arXiv:2605.16894v1 Announce Type: new Abstract: Reinforcement Learning (RL) uses rewards to guide learning, yet reward design is typically hand-crafted using heuristics that can be difficult to tune.

safetyarxiv-cs-ro
19 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios

DGX agent

arXiv:2605.16892v1 Announce Type: cross Abstract: Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification a

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Auditing Agent Harness Safety

DGX agent

arXiv:2605.14271v1 Announce Type: new Abstract: LLM agents increasingly run inside execution harnesses that dispatch tools, allocate resources, and route messages between specialized components. Howev

model-releasesarxiv-cs-cl
15 May 2026
Safety

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

DGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

safetyarxiv-cs-ai
14 May 2026
Model Releases

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

DGX agent

arXiv:2605.12015v1 Announce Type: cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools,

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries

DGX agent

arXiv:2605.01415v1 Announce Type: new Abstract: Recent AI systems compress the distance between capability growth and capability deployment. Earlier high-risk technologies were slowed by capital inten

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis

DGX agent

arXiv:2605.03441v1 Announce Type: cross Abstract: Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We s

model-releasesarxiv-cs-cl
6 May 2026
Local Ai

Edge AI for Automotive Vulnerable Road User Safety: Deployable Detection via Knowledge Distillation

DGX agent

arXiv:2604.26857v1 Announce Type: new Abstract: Deploying accurate object detection for Vulnerable Road User (VRU) safety on edge hardware requires balancing model capacity against computational const

local-aiarxiv-cs-cv
30 Apr 2026
Model Releases

Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training

DGX agent

arXiv:2510.20956v2 Announce Type: replace-cross Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaki

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

DGX agent

arXiv:2604.25102v1 Announce Type: new Abstract: Typographic prompt injection exploits vision language models' (VLMs) ability to read text rendered in images, posing a growing threat as VLMs power auto

model-releasesarxiv-cs-cv
29 Apr 2026
Local Ai

RADIANT-LLM: an Agentic Retrieval Augmented Generation Framework for Reliable Decision Support in Safety-Critical Nuclear Engineering

DGX agent

arXiv:2604.22755v1 Announce Type: cross Abstract: Reliable decision support in nuclear engineering requires traceable, domain-grounded knowledge retrieval, yet safety and risk analysis workflows remai

local-aiarxiv-cs-ai
28 Apr 2026
Safety

Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems

DGX agent

arXiv:2604.24153v1 Announce Type: new Abstract: Current AI systems increasingly operate in contexts where their outputs directly trigger real-world actions. Most existing approaches to AI safety, risk

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models

DGX agent

arXiv:2604.22156v1 Announce Type: cross Abstract: Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a comp

model-releasesarxiv-cs-cv
27 Apr 2026
Model Releases

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs

DGX agent

arXiv:2604.20945v1 Announce Type: cross Abstract: Effective safety auditing of large language models (LLMs) demands tools that go beyond black-box probing and systematically uncover vulnerabilities ro

model-releasesarxiv-cs-lg
24 Apr 2026
Model Releases

Owner-Harm: A Missing Threat Model for AI Agent Safety

DGX agent

arXiv:2604.18658v1 Announce Type: cross Abstract: Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a d

model-releasesarxiv-cs-ai
22 Apr 2026
Safety

A Real-Time Bike-Pedestrian Safety System with Wide-Angle Perception and Evaluation Testbed for Urban Intersections

DGX agent

arXiv:2604.17046v1 Announce Type: new Abstract: Collisions between cyclists and pedestrians at urban intersections remain a persistent source of injuries, yet few systems attempt real-time warnings to

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

Using large language models for embodied planning introduces systematic safety risks

DGX agent

arXiv:2604.18463v1 Announce Type: cross Abstract: Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe plann

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

DGX agent

arXiv:2505.10846v3 Announce Type: replace Abstract: This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its cor

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models

DGX agent

arXiv:2604.12447v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models inherit rich world knowledge from vision-language backbones and acquire executable skills via action demonstrations.

model-releasesarxiv-cs-ro
15 Apr 2026
Model Releases

Detecting Safety Violations Across Many Agent Traces

DGX agent

arXiv:2604.11806v1 Announce Type: new Abstract: To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex,

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Reliable and Real-Time Highway Trajectory Planning via Hybrid Learning-Optimization Frameworks

DGX agent

arXiv:2508.04436v2 Announce Type: replace Abstract: Autonomous highway driving involves high-speed safety risks due to limited reaction time, where rare but dangerous events may lead to severe consequ

safetyarxiv-cs-ro
14 Apr 2026
Model Releases

AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models

DGX agent

arXiv:2604.08867v1 Announce Type: cross Abstract: Audio has rapidly become a primary interface for foundation models, powering real-time voice assistants. Ensuring safety in audio systems is inherentl

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

PilotBench: A Benchmark for General Aviation Agents with Safety Constraints

DGX agent

arXiv:2604.08987v1 Announce Type: new Abstract: As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control

DGX agent

arXiv:2509.23960v2 Announce Type: replace-cross Abstract: Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Existing approaches based on multi-ag

safetyarxiv-cs-ai
7 Jul 2026
Safety

Revocable Learned State via Process Sidecars

DGX agent

arXiv:2606.30788v1 Announce Type: cross Abstract: Language models are often adapted in stages: a public skill phase, a private memory phase, and a later safety phase that learns to refuse outputs tied

safetyarxiv-cs-cl
1 Jul 2026
Safety

Mission-Level Runtime Assurance Framework for Autonomous Driving

DGX agent

arXiv:2606.06996v1 Announce Type: new Abstract: This paper studies runtime safety for autonomous driving when high-level driving commands become faulty or unreliable. Unlike conventional runtime-safet

safetyarxiv-cs-ro
8 Jun 2026
Safety

Perception with Guarantees: Certified Pose Estimation via Reachability Analysis

DGX agent

arXiv:2602.10032v2 Announce Type: replace Abstract: Agents in cyber-physical systems are increasingly entrusted with safety-critical tasks. Ensuring safety of these agents often requires localizing th

safetyarxiv-cs-cv
14 May 2026
Safety

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

DGX agent

arXiv:2509.12833v2 Announce Type: replace Abstract: Projection-based safety filters, which modify unsafe actions by mapping them to the closest safe alternative, are widely used to enforce safety cons

safetyarxiv-cs-lg
17 Apr 2026
Safety

Robust Real-Time Coordination of CAVs: A Distributed Optimization Framework under Uncertainty

DGX agent

arXiv:2508.21322v2 Announce Type: replace Abstract: Achieving both safety guarantees and real-time performance in cooperative vehicle coordination remains a fundamental challenge, particularly in dyna

safetyarxiv-cs-ro
14 Apr 2026
Model Releases

Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller

DGX agent

arXiv:2608.09653v1 Announce Type: cross Abstract: Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Certified Interpolation Oversampling: Per-Instance Safety Guarantees for Imbalanced Learning

DGX agent

arXiv:2501.15790v2 Announce Type: replace Abstract: Synthetic minority oversampling is typically designed and evaluated against a predictive objective, generating samples that improve downstream class

model-releasesarxiv-cs-lg
10 Aug 2026
Model Releases

StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

DGX agent

arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety

DGX agent

arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

DGX agent

arXiv:2607.28623v1 Announce Type: new Abstract: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body hum

model-releasesarxiv-cs-ro
31 Jul 2026
Model Releases

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

DGX agent

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Google reveals Gemini Robotics 2.0, promising improved dexterity and safety

DGX agent

Google announced Gemini Robotics 2.0 on July 30 2026, launching a family of three models that enhance robot dexterity, safety, and whole‑body intelligence for humanoid machinery. The publicly released

model-releasesars-technica
30 Jul 2026
Model Releases

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

DGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

DGX agent

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

model-releasesdair-ai--x
12 Jul 2026
Model Releases

Predicting LLM Safety Before Release by Simulating Deployment

DGX agent

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations

DGX agent

arXiv:2607.04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly co

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety resea…

DGX agent

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety research. So much more to do. We are 1% done. We've put together

model-releasesboris-cherny--x
6 Jul 2026
Model Releases

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

DGX agent

arXiv:2606.31644v1 Announce Type: new Abstract: As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behavio

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

DGX agent

arXiv:2606.28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown th

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repea…

DGX agent

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human

model-releasesopenai--x
26 Jun 2026
Model Releases

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation

DGX agent

arXiv:2606.25782v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs) in chatbots and everyday applications, companies increasingly need guardrails that are effe

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

DGX agent

arXiv:2606.18936v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature a

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

DGX agent

arXiv:2606.11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a

model-releasesarxiv-cs-lg
11 Jun 2026
Model Releases

Anthropic says Claude Fable 5 uses conservative safety classifiers that trigger a fallback to Claude Opus 4.8 in <5% of sessions, in areas like cybersecurity (Anthropic)

DGX agent

Anthropic: Anthropic says Claude Fable 5 uses conservative safety classifiers that trigger a fallback to Claude Opus 4.8 in <5% of sessions, in areas like cybersecurity — Today we're launching Claude

model-releasestechmeme
9 Jun 2026
← Previous
1…1314151617…297
Next →