AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
13 Apr 2026

Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers

Model ReleasesDGX agent

arXiv:2603.28013v3 Announce Type: replace-cross Abstract: Multi-agent LLM systems are entering production -- processing documents, managing workflows, acting on behalf of users -- yet their resilience

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.09452v1 Announce Type: cross Abstract: Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments

12 Apr 2026

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to se…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to see a debunking of their data, then we safety advocates should

10 Apr 2026

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Model ReleasesDGX agent

arXiv:2604.05172v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evalu

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

SafetyDGX agent

arXiv:2604.07754v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) raises significant ethical and safety concerns. While LLM alignment techniques are adopted to improve m

12 Aug 2026

Topological Feasibility Guarantees for Differentiable Predictive Control

SafetyDGX agent

arXiv:2608.10332v1 Announce Type: cross Abstract: Differentiable predictive control (DPC), a self-supervised learning approach for approximating explicit model predictive control (MPC) policies, offer

11 Aug 2026

InfoOps Bench: A live information operations safety benchmark

Model ReleasesDGX agent

arXiv:2607.28503v3 Announce Type: replace Abstract: In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted

7 Aug 2026

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

Model ReleasesDGX agent

arXiv:2608.06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.

6 Aug 2026

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

Model ReleasesDGX agent

arXiv:2608.05018v1 Announce Type: new Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as

5 Aug 2026

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

Model ReleasesDGX agent

arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f

4 Aug 2026

Grasp Execution Without a Planner: Configuration-Space Grasp Distance Fields with Certified Safety & Guaranteed Quality

Model ReleasesDGX agent

arXiv:2608.00600v1 Announce Type: new Abstract: Standard multifingered grasp execution architectures plan a collision-free trajectory to a selected grasp pose and track it with a feedback law. Executi

3 Aug 2026

White House invites AI companies to review its new AI safety framework

Model ReleasesDGX agent

Cybersecurity chiefs at the White House have reportedly finalized the outline of a forthcoming framework that will enable artificial intelligence companies to voluntarily submit their latest frontier

28 Jul 2026

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Model ReleasesDGX agent

arXiv:2607.24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. W

25 Jul 2026

Anthropic launches Claude Opus 5 with efficiency, safety improvements

Model ReleasesDGX agent

Anthropic PBC today rolled out a large language model called Claude Opus 5 to its chatbot service and developer platform. The company says the LLM approaches the output quality of its top-end Mythos 5

23 Jul 2026

DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning

Model ReleasesDGX agent

arXiv:2604.02694v2 Announce Type: replace-cross Abstract: The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safet

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

SafetyDGX agent

arXiv:2607.18325v1 Announce Type: new Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Visi

15 Jul 2026

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

SafetyDGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

7 Jul 2026

A Graph-Based Reinforcement Learning Approach with Frontier Potential Based Reward for Safe Cluttered Environment Exploration

SafetyDGX agent

arXiv:2504.11907v3 Announce Type: replace Abstract: Autonomous exploration of cluttered environments requires efficient exploration strategies that guarantee safety against potential collisions with u

Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

SafetyDGX agent

arXiv:2508.10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and

1 Jul 2026

Anthropic launches Claude Sonnet 5 AI model with coding, safety upgrades as Fable and Mythos controls lifted

Model ReleasesDGX agent

Anthropic PBC today debuted Claude Sonnet 5, a midrange large language model that outperforms its predecessor in several areas. The LLM will be the default option in the consumer tiers of the company’

30 Jun 2026

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2606.22528v2 Announce Type: replace Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show t

Lateral String Stability for Vehicle Platoons

SafetyDGX agent

arXiv:2606.29677v1 Announce Type: new Abstract: Connected and automated vehicle (CAV) platooning promises gains in energy efficiency and traffic throughput and, most critically, in safety. These safet

29 Jun 2026

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

SafetyDGX agent

arXiv:2606.28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively

26 Jun 2026

Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation

Local AiDGX agent

arXiv:2606.26686v1 Announce Type: new Abstract: In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict. This design follo

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

SafetyDGX agent

arXiv:2606.26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient auton

23 Jun 2026

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

SafetyDGX agent

arXiv:2606.20698v1 Announce Type: new Abstract: Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However,

11 Jun 2026

Runtime Enforcement of Hybrid System Properties

SafetyDGX agent

arXiv:2606.12022v1 Announce Type: cross Abstract: Runtime enforcement has emerged as a promising approach for ensuring the safety of autonomous and cyber-physical systems operating in uncertain and dy

8 Jun 2026

An Integrated Roadside Sensing and Communication Framework for Vulnerable Road User Safety at Signalized Intersections

Model ReleasesDGX agent

arXiv:2606.07016v1 Announce Type: cross Abstract: Vulnerable road users (VRUs) account for approximately half of urban traffic deaths globally, with intersections concentrating a disproportionate shar

4 Jun 2026

Explainably Safe Reinforcement Learning

SafetyDGX agent

arXiv:2606.04634v1 Announce Type: new Abstract: Trust in a decision-making system requires both safety guarantees and the ability to interpret and understand its behavior. This is particularly importa

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

SafetyDGX agent

arXiv:2511.07107v3 Announce Type: replace Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address im

3 Jun 2026

MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for Medical AI Safety

Model ReleasesDGX agent

arXiv:2606.02630v1 Announce Type: cross Abstract: Patient-facing medical chatbots are commonly evaluated on single-turn prompts, yet real users push back after refusals, add urgency, and invoke author

2 Jun 2026

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

Model ReleasesDGX agent

arXiv:2606.02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

Model ReleasesDGX agent

arXiv:2606.00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode colla

Robust Shielding for Safe Reinforcement Learning

SafetyDGX agent

arXiv:2606.00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs). However, exist

27 May 2026

Device Context Protocol: A Compact, Safety-First Architecture for LLM-Driven Control of Constrained Devices

Model ReleasesDGX agent

arXiv:2605.26159v1 Announce Type: cross Abstract: Large language models are increasingly used as orchestrators of external tools via the Model Context Protocol (MCP), but MCP is built for software ser

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

Model ReleasesDGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

20 May 2026

Guiding Neuro-Symbolic Scenario Generation with Spatio-Temporal Logic

SafetyDGX agent

arXiv:2605.19038v1 Announce Type: cross Abstract: The rapid advancement of autonomous driving (AD) technologies has outpaced the development of robust safety evaluation methods. Conventional testing r

19 May 2026

LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails

SafetyDGX agent

arXiv:2605.17329v1 Announce Type: cross Abstract: Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, s

15 May 2026

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

Local AiDGX agent

arXiv:2605.14454v1 Announce Type: cross Abstract: As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line

Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks

Model ReleasesDGX agent

arXiv:2605.14604v1 Announce Type: new Abstract: This position paper argues that effective tutoring requires corrective friction: surfacing misconceptions and challenging them supportively to drive con

13 May 2026

Metaphor Is Not All Attention Needs

SafetyDGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

12 May 2026

Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization

SafetyDGX agent

arXiv:2409.10310v3 Announce Type: replace Abstract: Ensuring safety and driving consistency is a significant challenge for autonomous vehicles operating in partially observed environments. This work i

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

Local AiDGX agent

arXiv:2605.08878v1 Announce Type: cross Abstract: Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representa

11 May 2026

GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access

Model ReleasesDGX agent

Executive Summary Since our February 2026 report on AI-related threat activity, Google Threat Intelligence Group (GTIG) has continued to track a maturing transition from nascent AI-enabled operations

5 May 2026

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

SafetyDGX agent

arXiv:2601.19487v2 Announce Type: replace Abstract: Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector

4 May 2026

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

Model ReleasesDGX agent

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal

1 May 2026

Connected Dependability Cage: Run-Time Function and Anomaly Monitoring for the Development and Operation of Safe Automated Vehicles

SafetyDGX agent

arXiv:2604.27728v1 Announce Type: new Abstract: The advancement of automated vehicles introduces complex safety challenges, particularly in dynamic and unpredictable environments where AI-enabled perc

30 Apr 2026

Culturally Aware GenAI Risks for Youth: Perspectives from Youth, Parents, and Teachers in a Non-Western Context

SafetyDGX agent

arXiv:2604.26494v1 Announce Type: cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youth's safety i

27 Apr 2026

Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

SafetyDGX agent

arXiv:2604.22260v1 Announce Type: cross Abstract: Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While rece

23 Apr 2026

Atomic Decision Boundaries: A Structural Requirement for Guaranteeing Execution-Time Admissibility in Autonomous Systems

SafetyDGX agent

arXiv:2604.17511v2 Announce Type: replace-cross Abstract: Autonomous systems increasingly execute actions that directly modify shared state, creating an urgent need for precise control over which tran

Stochastic Barrier Certificates in the Presence of Dynamic Obstacles

SafetyDGX agent

arXiv:2604.20208v1 Announce Type: new Abstract: Safety of stochastic dynamic systems in environments with dynamic obstacles is studied in this paper through the lens of stochastic barrier functions. W

21 Apr 2026

Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2512.16055v2 Announce Type: replace Abstract: Safety-critical corner cases, difficult to collect in the real world, are crucial for evaluating end-to-end autonomous driving. Adversarial interact

J-PARSE: Jacobian-based Projection Algorithm for Resolving Singularities Effectively in Inverse Kinematic Control of Serial Manipulators

SafetyDGX agent

arXiv:2505.00306v5 Announce Type: replace Abstract: J-PARSE is an algorithm for smooth first-order inverse kinematic control of a serial manipulator near kinematic singularities. The commanded end-eff

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

SafetyDGX agent

arXiv:2604.17769v1 Announce Type: new Abstract: Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-e

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

SafetyDGX agent

arXiv:2604.16606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

SafetyDGX agent

arXiv:2601.04740v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risk

20 Apr 2026

FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

SafetyDGX agent

arXiv:2604.15488v1 Announce Type: cross Abstract: Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offer

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

Model ReleasesDGX agent

arXiv:2505.24672v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploit

16 Apr 2026

Drowsiness-Aware Adaptive Autonomous Braking System based on Deep Reinforcement Learning for Enhanced Road Safety

Model ReleasesDGX agent

arXiv:2604.13878v1 Announce Type: new Abstract: Driver drowsiness significantly impairs the ability to accurately judge safe braking distances and is estimated to contribute to 10%-20% of road acciden

15 Apr 2026

ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack

SafetyDGX agent

arXiv:2509.25843v2 Announce Type: replace Abstract: Large language models (LLMs), despite being safety-aligned, exhibit brittle refusal behaviors that can be circumvented by simple linguistic changes.

← Previous
1…1213141516…238
Next →