AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
10 Jun 2026

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. …

SafetyDGX agent

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't d

9 Jun 2026

CARE: A Conformal Safety Layer for Medical Summarization

SafetyDGX agent

arXiv:2606.08969v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce

fable’s safety guardrails broken within an hour. raise your hand if you are surprised.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

fable’s safety guardrails broken within an hour. raise your hand if you are surprised. We tested Anthropic’s new @claudeai Fable 5. It did not fail like an ordinary jailbreak. It failed more quietly.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

Model ReleasesDGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

SafetyDGX agent

arXiv:2606.08860v1 Announce Type: new Abstract: Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing from digital maps, creating safety risks f

5 Jun 2026

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

Model ReleasesDGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Ye…

SafetyDGX agent

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Yet most of the field kept licking the boot of the companies a

Safety officials finally have a good idea of what a big rocket explosion can do

SafetyDGX agent

Safety officials gained detailed data on the impact of large rocket explosions when Blue Origin's New Glenn rocket experienced an anomaly during a hot-fire test at Cape Canaveral on May 28, causing a

Sources and docs detail defense tech startup Shield AI's struggles to overcome years of technical hitches and safety concerns with its V-BAT autonomous drone (David Jeans/Reuters)

SafetyDGX agent

David Jeans / Reuters: Sources and docs detail defense tech startup Shield AI's struggles to overcome years of technical hitches and safety concerns with its V-BAT autonomous drone — A year ago, Ryan

4 Jun 2026

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

SafetyDGX agent

arXiv:2602.06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications,

Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

Model ReleasesDGX agent

arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5

2 Jun 2026

Advancing youth safety and opportunity through global leadership

SafetyDGX agent

OpenAI outlines initiatives and commitments to enhance youth safety while creating educational and economic opportunities for young people globally. The effort emphasizes OpenAI's role in responsible

SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.00773v1 Announce Type: new Abstract: Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety-relevant tr

30 May 2026

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the o…

SafetyDGX agent

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the open on @huggingface, so researchers everywhere can scrutiniz

29 May 2026

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

SafetyDGX agent

arXiv:2605.29340v1 Announce Type: new Abstract: In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual an

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

Model ReleasesDGX agent

arXiv:2605.29801v1 Announce Type: new Abstract: Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhi

Former Tesla data labelers say FSD relies on laborious mapping for hazards; crash data analysis shows Tesla exaggerates FSD's safety via flawed methodology (Reuters)

SafetyDGX agent

Reuters: Former Tesla data labelers say FSD relies on laborious mapping for hazards; crash data analysis shows Tesla exaggerates FSD's safety via flawed methodology — Tesla says its Full Self-Driving

28 May 2026

Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.28552v1 Announce Type: new Abstract: As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle ty

27 May 2026

AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

Model ReleasesDGX agent

arXiv:2605.26954v1 Announce Type: new Abstract: Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved.

Illinois Legislature passes SB 315, a bill requiring annual independent third-party safety audits of leading AI companies; the bill heads to the governor's desk (Jared Perlo/NBC News)

SafetyDGX agent

Jared Perlo / NBC News: Illinois Legislature passes SB 315, a bill requiring annual independent third-party safety audits of leading AI companies; the bill heads to the governor's desk — The measure,

Position: AI Safety Requires Effective Controllability

Model ReleasesDGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

SafetyDGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

26 May 2026

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_!

SafetyDGX agent

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_! Is AI development progressing too quickly? @business' @shiringh

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

Model ReleasesDGX agent

arXiv:2605.24154v1 Announce Type: new Abstract: Current safety alignment of foundation models largely follows a one-size-fits-all paradigm, applying the same refusal policy across users and contexts.

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

Model ReleasesDGX agent

arXiv:2601.21094v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safe

23 May 2026

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

SafetyDGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

Q&A with Sundar Pichai on the future of Google Search, Google's place in the AI race, public skepticism toward AI, AI agents, AI safety, TPUs, and more (New York Times)

SafetyDGX agent

New York Times: Q&A with Sundar Pichai on the future of Google Search, Google's place in the AI race, public skepticism toward AI, AI agents, AI safety, TPUs, and more — After a busy Google I/O, the c

22 May 2026

CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety

SafetyDGX agent

arXiv:2605.21609v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating information seeking, advice, and emotionally sensit

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

SafetyDGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks t…

SafetyDGX agent

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks things are fine readiness wise or incentive wise etc. https:/

21 May 2026

Safety-Critical Control for Smoothed Implicit Contact Dynamics

Model ReleasesDGX agent

arXiv:2605.21138v1 Announce Type: new Abstract: Smoothed implicit contact dynamics enables gradient-based planning and control for contact-rich tasks without predefined mode sequences. However, safety

20 May 2026

OpenAI's Chris Lehane says he is pursuing 'reverse federalism', lobbying blue states to pass AI safety laws and create a de facto US standard, as DC dithers (Brendan Bordelon/Politico)

SafetyDGX agent

Brendan Bordelon / Politico: OpenAI's Chris Lehane says he is pursuing “reverse federalism”, lobbying blue states to pass AI safety laws and create a de facto US standard, as DC dithers — OpenAI's eff

19 May 2026

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

Model ReleasesDGX agent

arXiv:2605.17971v1 Announce Type: cross Abstract: Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuri

18 May 2026

Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks

Model ReleasesDGX agent

arXiv:2603.04459v3 Announce Type: replace-cross Abstract: The rapid expansion of research in LLM safety presents challenges in tracking advancements, making benchmarks important evaluation infrastruct

parallelcbf: A composable safety-filter and auditability framework for tensor-parallel reinforcement learning

SafetyDGX agent

arXiv:2605.15509v1 Announce Type: new Abstract: While Isaac Lab provides massive parallel UAV simulation, OmniSafe and safe-control-gym provide constrained-RL benchmarks, and CBFKit provides control-b

15 May 2026

OpenAI's disavowal of a liability shield in Illinois SB 3444 bill and endorsement of a stronger SB 315 suggest it is open to meaningful AI safety legislation (Transformer)

SafetyDGX agent

Transformer: OpenAI's disavowal of a liability shield in Illinois SB 3444 bill and endorsement of a stronger SB 315 suggest it is open to meaningful AI safety legislation — Transformer Weekly: US-Chin

11 May 2026

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

SafetyDGX agent

arXiv:2605.07250v1 Announce Type: cross Abstract: Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we i

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

Model ReleasesDGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

8 May 2026

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can fi…

SafetyDGX agent

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can find @redwood_ai's report here: https://blog.redwoodresearch.o

7 May 2026

Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content

SafetyDGX agent

arXiv:2605.04085v1 Announce Type: cross Abstract: Objectives: Large language models (LLMs) are increasingly used for clinical text summarization, yet structured methods to assess associated patient sa

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founde…

SafetyDGX agent

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founders of Anthropic such as @DarioAmodei and @jackclarkSF may ha

Meta challenges Ofcom in UK High Court over the Online Safety Act, which calculates levies based on global, not UK, revenue, in a case scheduled for October (Sam Tobin/Reuters)

SafetyDGX agent

Sam Tobin / Reuters: Meta challenges Ofcom in UK High Court over the Online Safety Act, which calculates levies based on global, not UK, revenue, in a case scheduled for October — Facebook and Instagr

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments

SafetyDGX agent

arXiv:2508.04204v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have demonstrated impressive performance in reasoning-intensive tasks, but they remain vulnerable to harmful content g

6 May 2026

Musk v. Altman: Mira Murati testifies that Sam Altman lied to her about the safety standards for a new OpenAI model and that he made her work more difficult (Jay Peters/The Verge)

SafetyDGX agent

Jay Peters / The Verge: Musk v. Altman: Mira Murati testifies that Sam Altman lied to her about the safety standards for a new OpenAI model and that he made her work more difficult — OpenAI's former C

5 May 2026

Improving Model Safety by Targeted Error Correction

SafetyDGX agent

arXiv:2605.02544v1 Announce Type: cross Abstract: The widespread adoption of machine learning in critical applications demands techniques to mitigate high-consequence errors. Our method utilizes a dua

4 May 2026

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

Model ReleasesDGX agent

arXiv:2605.00689v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments

New Mexico child safety trial: New Mexico asks a judge to declare Meta a public nuisance and to order it to pay $3.7B and overhaul its apps to protect children (Diana Novak Jones/Reuters)

SafetyDGX agent

Diana Novak Jones / Reuters: New Mexico child safety trial: New Mexico asks a judge to declare Meta a public nuisance and to order it to pay $3.7B and overhaul its apps to protect children — The U.S.

3 May 2026

A profile of BlackBerry's QNX division, whose operating system controls safety features in 275M cars and accounts for half of BlackBerry's revenue (Ben Cohen/Wall Street Journal)

SafetyDGX agent

Ben Cohen / Wall Street Journal: A profile of BlackBerry's QNX division, whose operating system controls safety features in 275M cars and accounts for half of BlackBerry's revenue — John Wall has spen

1 May 2026

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

Model ReleasesDGX agent

arXiv:2604.26959v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into patient-facing healthcare systems offers significant potential to improve access to medical information.

GAVEL: Towards Rule-Based Safety Through Activation Monitoring

SafetyDGX agent

arXiv:2601.19768v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be appare

30 Apr 2026

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely un…

SafetyDGX agent

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely unregulated. You just can’t. - Gary The judge just instructed

Test-Time Safety Alignment

SafetyDGX agent

arXiv:2604.26167v1 Announce Type: cross Abstract: Recent work has shown that a model's input word embeddings can serve as effective control variables for steering its behavior toward outputs that sati

28 Apr 2026

A Lightweight Explainable Guardrail for Prompt Safety

SafetyDGX agent

arXiv:2602.15853v2 Announce Type: replace-cross Abstract: We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly l

Does Machine Unlearning Preserve Clinical Safety? A Risk Analysis for Medical Image Classification

SafetyDGX agent

arXiv:2604.23854v1 Announce Type: new Abstract: The application of Deep Learning in medical diagnosis must balance patient safety with compliance with data protection regulations. Machine Unlearning e

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models

Model ReleasesDGX agent

arXiv:2511.08484v2 Announce Type: replace Abstract: We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilitie

Time-Series Forecasting in Safety-Critical Environments: An EU-AI-Act-Compliant Open-Source Package / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein KI-VO-konformes Open-Source-Paket

SafetyDGX agent

arXiv:2604.23859v1 Announce Type: new Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical envir

22 Apr 2026

Australia's eSafety Commissioner issues transparency notices to Roblox, Microsoft's Minecraft, and other online gaming platforms to detail child safety measures (Renju Jose/Reuters)

SafetyDGX agent

Renju Jose / Reuters: Australia's eSafety Commissioner issues transparency notices to Roblox, Microsoft's Minecraft, and other online gaming platforms to detail child safety measures — Australia's int

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

21 Apr 2026

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinct…

SafetyDGX agent

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinction risk. We're not. It's not that complicated: AI smarter t

17 Apr 2026

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs

SafetyDGX agent

arXiv:2509.05367v4 Announce Type: replace-cross Abstract: Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classificati

← Previous
1…678910…238
Next →