AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
10 Jun 2026

For Robotaxis, Safety Must Be Built In, Not Bolted On

SafetyDGX agent

A car pulls up to the curb. The app says, “Your ride is here.” No one’s in the driver’s seat. For people who live in one of the dozens of cities now hosting robotaxi services, this is already a realit

9 Jun 2026

Claude Fable 5 and new AI safety fables

Model ReleasesDGX agent

This article discusses Claude Fable 5, likely exploring Anthropic's latest version of their AI model and examining new fables or narratives related to AI safety concepts. The piece probably analyzes h

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.07929v1 Announce Type: new Abstract: Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we p

Anthropic says Claude Fable 5 uses conservative safety classifiers that trigger a fallback to Claude Opus 4.8 in <5% of sessions, in areas like cybersecurity (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic says Claude Fable 5 uses conservative safety classifiers that trigger a fallback to Claude Opus 4.8 in <5% of sessions, in areas like cybersecurity — Today we're launching Claude

8 Jun 2026

An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?

Model ReleasesDGX agent

arXiv:2605.25806v2 Announce Type: replace Abstract: Women's safety and security are paramount for a modern society. Crimes against women occur in daylight as well as in low-light conditions. Often, su

Mission-Level Runtime Assurance Framework for Autonomous Driving

SafetyDGX agent

arXiv:2606.06996v1 Announce Type: new Abstract: This paper studies runtime safety for autonomous driving when high-level driving commands become faulty or unreliable. Unlike conventional runtime-safet

2 Jun 2026

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

Model ReleasesDGX agent

arXiv:2606.00813v1 Announce Type: cross Abstract: Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with qua

Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

SafetyDGX agent

arXiv:2606.02444v1 Announce Type: new Abstract: Recent evidence shows that people with eating disorders (EDs) are increasingly seeking guidance, advice, and emotional support from Large Language Model

InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation

SafetyDGX agent

arXiv:2512.10234v2 Announce Type: replace-cross Abstract: Even LLMs that appear safe during evaluation can still produce harmful responses in deployment. Because stochastic sampling yields different r

Market-Based Replanning for Safety-Critical UAV Swarms in Search and Rescue Missions

SafetyDGX agent

arXiv:2606.01970v1 Announce Type: new Abstract: Reliable autonomous UAV swarms in Search and Rescue (SAR) missions require fault-tolerant coordination capable of sustaining operations despite agent de

29 May 2026

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

28 May 2026

No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees

SafetyDGX agent

arXiv:2603.08761v2 Announce Type: replace-cross Abstract: We argue that formal certification of AI alignment over open-ended or unbounded input domains is impossible under standard assumptions in comp

25 May 2026

Safe Reinforcement Learning with Preference-based Constraint Inference

SafetyDGX agent

arXiv:2603.23565v2 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) is a standard paradigm for safety-critical decision making. However, real-world safety constraints can be com

22 May 2026

Boundary-targeted Membership Inference Attacks on Safety Classifiers

Local AiDGX agent

arXiv:2605.22373v1 Announce Type: cross Abstract: Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with

19 May 2026

Beyond Safety Filtering: Control Barrier Function-Informed Reinforcement Learning for Connected and Automated Vehicles

SafetyDGX agent

arXiv:2605.16894v1 Announce Type: new Abstract: Reinforcement Learning (RL) uses rewards to guide learning, yet reward design is typically hand-crafted using heuristics that can be difficult to tune.

DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios

Model ReleasesDGX agent

arXiv:2605.16892v1 Announce Type: cross Abstract: Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification a

15 May 2026

Auditing Agent Harness Safety

Model ReleasesDGX agent

arXiv:2605.14271v1 Announce Type: new Abstract: LLM agents increasingly run inside execution harnesses that dispatch tools, allocate resources, and route messages between specialized components. Howev

14 May 2026

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

SafetyDGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

Perception with Guarantees: Certified Pose Estimation via Reachability Analysis

SafetyDGX agent

arXiv:2602.10032v2 Announce Type: replace Abstract: Agents in cyber-physical systems are increasingly entrusted with safety-critical tasks. Ensuring safety of these agents often requires localizing th

13 May 2026

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

Model ReleasesDGX agent

arXiv:2605.12015v1 Announce Type: cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools,

6 May 2026

AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries

Model ReleasesDGX agent

arXiv:2605.01415v1 Announce Type: new Abstract: Recent AI systems compress the distance between capability growth and capability deployment. Earlier high-risk technologies were slowed by capital inten

Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis

Model ReleasesDGX agent

arXiv:2605.03441v1 Announce Type: cross Abstract: Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We s

30 Apr 2026

Edge AI for Automotive Vulnerable Road User Safety: Deployable Detection via Knowledge Distillation

Local AiDGX agent

arXiv:2604.26857v1 Announce Type: new Abstract: Deploying accurate object detection for Vulnerable Road User (VRU) safety on edge hardware requires balancing model capacity against computational const

Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training

Model ReleasesDGX agent

arXiv:2510.20956v2 Announce Type: replace-cross Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaki

29 Apr 2026

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

Model ReleasesDGX agent

arXiv:2604.25102v1 Announce Type: new Abstract: Typographic prompt injection exploits vision language models' (VLMs) ability to read text rendered in images, posing a growing threat as VLMs power auto

28 Apr 2026

RADIANT-LLM: an Agentic Retrieval Augmented Generation Framework for Reliable Decision Support in Safety-Critical Nuclear Engineering

Local AiDGX agent

arXiv:2604.22755v1 Announce Type: cross Abstract: Reliable decision support in nuclear engineering requires traceable, domain-grounded knowledge retrieval, yet safety and risk analysis workflows remai

Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems

SafetyDGX agent

arXiv:2604.24153v1 Announce Type: new Abstract: Current AI systems increasingly operate in contexts where their outputs directly trigger real-world actions. Most existing approaches to AI safety, risk

27 Apr 2026

Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.22156v1 Announce Type: cross Abstract: Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a comp

24 Apr 2026

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs

Model ReleasesDGX agent

arXiv:2604.20945v1 Announce Type: cross Abstract: Effective safety auditing of large language models (LLMs) demands tools that go beyond black-box probing and systematically uncover vulnerabilities ro

22 Apr 2026

Owner-Harm: A Missing Threat Model for AI Agent Safety

Model ReleasesDGX agent

arXiv:2604.18658v1 Announce Type: cross Abstract: Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a d

21 Apr 2026

A Real-Time Bike-Pedestrian Safety System with Wide-Angle Perception and Evaluation Testbed for Urban Intersections

SafetyDGX agent

arXiv:2604.17046v1 Announce Type: new Abstract: Collisions between cyclists and pedestrians at urban intersections remain a persistent source of injuries, yet few systems attempt real-time warnings to

Using large language models for embodied planning introduces systematic safety risks

Model ReleasesDGX agent

arXiv:2604.18463v1 Announce Type: cross Abstract: Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe plann

17 Apr 2026

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

Model ReleasesDGX agent

arXiv:2505.10846v3 Announce Type: replace Abstract: This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its cor

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

SafetyDGX agent

arXiv:2509.12833v2 Announce Type: replace Abstract: Projection-based safety filters, which modify unsafe actions by mapping them to the closest safe alternative, are widely used to enforce safety cons

15 Apr 2026

HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2604.12447v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models inherit rich world knowledge from vision-language backbones and acquire executable skills via action demonstrations.

14 Apr 2026

Detecting Safety Violations Across Many Agent Traces

Model ReleasesDGX agent

arXiv:2604.11806v1 Announce Type: new Abstract: To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex,

Reliable and Real-Time Highway Trajectory Planning via Hybrid Learning-Optimization Frameworks

SafetyDGX agent

arXiv:2508.04436v2 Announce Type: replace Abstract: Autonomous highway driving involves high-speed safety risks due to limited reaction time, where rare but dangerous events may lead to severe consequ

Robust Real-Time Coordination of CAVs: A Distributed Optimization Framework under Uncertainty

SafetyDGX agent

arXiv:2508.21322v2 Announce Type: replace Abstract: Achieving both safety guarantees and real-time performance in cooperative vehicle coordination remains a fundamental challenge, particularly in dyna

13 Apr 2026

AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models

Model ReleasesDGX agent

arXiv:2604.08867v1 Announce Type: cross Abstract: Audio has rapidly become a primary interface for foundation models, powering real-time voice assistants. Ensuring safety in audio systems is inherentl

PilotBench: A Benchmark for General Aviation Agents with Safety Constraints

Model ReleasesDGX agent

arXiv:2604.08987v1 Announce Type: new Abstract: As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained

7 Jul 2026

MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control

SafetyDGX agent

arXiv:2509.23960v2 Announce Type: replace-cross Abstract: Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Existing approaches based on multi-ag

Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations

Model ReleasesDGX agent

arXiv:2607.04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly co

1 Jul 2026

Revocable Learned State via Process Sidecars

SafetyDGX agent

arXiv:2606.30788v1 Announce Type: cross Abstract: Language models are often adapted in stages: a public skill phase, a private memory phase, and a later safety phase that learns to refuse outputs tied

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

Model ReleasesDGX agent

arXiv:2606.31644v1 Announce Type: new Abstract: As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behavio

11 Aug 2026

Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller

Model ReleasesDGX agent

arXiv:2608.09653v1 Announce Type: cross Abstract: Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since

10 Aug 2026

Certified Interpolation Oversampling: Per-Instance Safety Guarantees for Imbalanced Learning

Model ReleasesDGX agent

arXiv:2501.15790v2 Announce Type: replace Abstract: Synthetic minority oversampling is typically designed and evaluated against a predictive objective, generating samples that improve downstream class

StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

Model ReleasesDGX agent

arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as

4 Aug 2026

Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety

Model ReleasesDGX agent

arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc

31 Jul 2026

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

Model ReleasesDGX agent

arXiv:2607.28623v1 Announce Type: new Abstract: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body hum

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

Model ReleasesDGX agent

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

30 Jul 2026

Google reveals Gemini Robotics 2.0, promising improved dexterity and safety

Model ReleasesDGX agent

Google announced Gemini Robotics 2.0 on July 30 2026, launching a family of three models that enhance robot dexterity, safety, and whole‑body intelligence for humanoid machinery. The publicly released

28 Jul 2026

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

Model ReleasesDGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

12 Jul 2026

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

Model ReleasesDGX agent

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

9 Jul 2026

Predicting LLM Safety Before Release by Simulating Deployment

Model ReleasesDGX agent

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about

6 Jul 2026

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety resea…

Model ReleasesDGX agent

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety research. So much more to do. We are 1% done. We've put together

30 Jun 2026

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown th

26 Jun 2026

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repea…

Model ReleasesDGX agent

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human

25 Jun 2026

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation

Model ReleasesDGX agent

arXiv:2606.25782v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs) in chatbots and everyday applications, companies increasingly need guardrails that are effe

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Model ReleasesDGX agent

arXiv:2606.18936v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature a

11 Jun 2026

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

Model ReleasesDGX agent

arXiv:2606.11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a

← Previous
1…1011121314…238
Next →