AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Model Releases

SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models

DGX agent

arXiv:2606.00773v1 Announce Type: new Abstract: Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety-relevant tr

model-releasesarxiv-cs-ro
2 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the o…

DGX agent

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the open on @huggingface, so researchers everywhere can scrutiniz

safetyclem-delangue--x
30 May 2026
Safety

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

DGX agent

arXiv:2605.29340v1 Announce Type: new Abstract: In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual an

safetyarxiv-cs-cl
29 May 2026
Model Releases

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

DGX agent

arXiv:2605.29801v1 Announce Type: new Abstract: Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhi

model-releasesarxiv-cs-ai
29 May 2026
Safety

Former Tesla data labelers say FSD relies on laborious mapping for hazards; crash data analysis shows Tesla exaggerates FSD's safety via flawed methodology (Reuters)

DGX agent

Reuters: Former Tesla data labelers say FSD relies on laborious mapping for hazards; crash data analysis shows Tesla exaggerates FSD's safety via flawed methodology — Tesla says its Full Self-Driving

safetytechmeme
29 May 2026
Safety

Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning

DGX agent

arXiv:2605.28552v1 Announce Type: new Abstract: As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle ty

safetyarxiv-cs-ai
28 May 2026
Model Releases

AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

DGX agent

arXiv:2605.26954v1 Announce Type: new Abstract: Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved.

model-releasesarxiv-cs-cl
27 May 2026
Safety

Illinois Legislature passes SB 315, a bill requiring annual independent third-party safety audits of leading AI companies; the bill heads to the governor's desk (Jared Perlo/NBC News)

DGX agent

Jared Perlo / NBC News: Illinois Legislature passes SB 315, a bill requiring annual independent third-party safety audits of leading AI companies; the bill heads to the governor's desk — The measure,

safetytechmeme
27 May 2026
Model Releases

Position: AI Safety Requires Effective Controllability

DGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

model-releasesarxiv-cs-ai
27 May 2026
Safety

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

DGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

safetyarxiv-cs-ai
27 May 2026
Safety

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_!

DGX agent

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_! Is AI development progressing too quickly? @business' @shiringh

safetyyoshua-bengio--x
26 May 2026
Model Releases

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

DGX agent

arXiv:2605.24154v1 Announce Type: new Abstract: Current safety alignment of foundation models largely follows a one-size-fits-all paradigm, applying the same refusal policy across users and contexts.

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

DGX agent

arXiv:2601.21094v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safe

model-releasesarxiv-cs-ai
26 May 2026
Safety

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

DGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

safetygary-marcus--x
23 May 2026
Safety

Q&A with Sundar Pichai on the future of Google Search, Google's place in the AI race, public skepticism toward AI, AI agents, AI safety, TPUs, and more (New York Times)

DGX agent

New York Times: Q&A with Sundar Pichai on the future of Google Search, Google's place in the AI race, public skepticism toward AI, AI agents, AI safety, TPUs, and more — After a busy Google I/O, the c

safetytechmeme
23 May 2026
Safety

CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety

DGX agent

arXiv:2605.21609v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating information seeking, advice, and emotionally sensit

safetyarxiv-cs-cl
22 May 2026
Safety

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

DGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

safetyarxiv-cs-cv
22 May 2026
Safety

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks t…

DGX agent

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks things are fine readiness wise or incentive wise etc. https:/

safetygary-marcus--x
22 May 2026
Model Releases

Safety-Critical Control for Smoothed Implicit Contact Dynamics

DGX agent

arXiv:2605.21138v1 Announce Type: new Abstract: Smoothed implicit contact dynamics enables gradient-based planning and control for contact-rich tasks without predefined mode sequences. However, safety

model-releasesarxiv-cs-ro
21 May 2026
Safety

OpenAI's Chris Lehane says he is pursuing 'reverse federalism', lobbying blue states to pass AI safety laws and create a de facto US standard, as DC dithers (Brendan Bordelon/Politico)

DGX agent

Brendan Bordelon / Politico: OpenAI's Chris Lehane says he is pursuing “reverse federalism”, lobbying blue states to pass AI safety laws and create a de facto US standard, as DC dithers — OpenAI's eff

safetytechmeme
20 May 2026
Model Releases

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

DGX agent

arXiv:2605.17971v1 Announce Type: cross Abstract: Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuri

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks

DGX agent

arXiv:2603.04459v3 Announce Type: replace-cross Abstract: The rapid expansion of research in LLM safety presents challenges in tracking advancements, making benchmarks important evaluation infrastruct

model-releasesarxiv-cs-ai
18 May 2026
Safety

parallelcbf: A composable safety-filter and auditability framework for tensor-parallel reinforcement learning

DGX agent

arXiv:2605.15509v1 Announce Type: new Abstract: While Isaac Lab provides massive parallel UAV simulation, OmniSafe and safe-control-gym provide constrained-RL benchmarks, and CBFKit provides control-b

safetyarxiv-cs-lg
18 May 2026
Safety

OpenAI's disavowal of a liability shield in Illinois SB 3444 bill and endorsement of a stronger SB 315 suggest it is open to meaningful AI safety legislation (Transformer)

DGX agent

Transformer: OpenAI's disavowal of a liability shield in Illinois SB 3444 bill and endorsement of a stronger SB 315 suggest it is open to meaningful AI safety legislation — Transformer Weekly: US-Chin

safetytechmeme
15 May 2026
Safety

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

DGX agent

arXiv:2605.07250v1 Announce Type: cross Abstract: Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we i

safetyarxiv-cs-ai
11 May 2026
Model Releases

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

DGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

model-releasesarxiv-cs-ai
11 May 2026
Safety

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can fi…

DGX agent

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can find @redwood_ai's report here: https://blog.redwoodresearch.o

safetyopenai--x
8 May 2026
Safety

Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content

DGX agent

arXiv:2605.04085v1 Announce Type: cross Abstract: Objectives: Large language models (LLMs) are increasingly used for clinical text summarization, yet structured methods to assess associated patient sa

safetyarxiv-cs-cl
7 May 2026
Safety

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founde…

DGX agent

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founders of Anthropic such as @DarioAmodei and @jackclarkSF may ha

safetygary-marcus--x
7 May 2026
Safety

Meta challenges Ofcom in UK High Court over the Online Safety Act, which calculates levies based on global, not UK, revenue, in a case scheduled for October (Sam Tobin/Reuters)

DGX agent

Sam Tobin / Reuters: Meta challenges Ofcom in UK High Court over the Online Safety Act, which calculates levies based on global, not UK, revenue, in a case scheduled for October — Facebook and Instagr

safetytechmeme
7 May 2026
Safety

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments

DGX agent

arXiv:2508.04204v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have demonstrated impressive performance in reasoning-intensive tasks, but they remain vulnerable to harmful content g

safetyarxiv-cs-cl
7 May 2026
Safety

Musk v. Altman: Mira Murati testifies that Sam Altman lied to her about the safety standards for a new OpenAI model and that he made her work more difficult (Jay Peters/The Verge)

DGX agent

Jay Peters / The Verge: Musk v. Altman: Mira Murati testifies that Sam Altman lied to her about the safety standards for a new OpenAI model and that he made her work more difficult — OpenAI's former C

safetytechmeme
6 May 2026
Safety

Improving Model Safety by Targeted Error Correction

DGX agent

arXiv:2605.02544v1 Announce Type: cross Abstract: The widespread adoption of machine learning in critical applications demands techniques to mitigate high-consequence errors. Our method utilizes a dua

safetyarxiv-cs-cv
5 May 2026
Model Releases

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

DGX agent

arXiv:2605.00689v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments

model-releasesarxiv-cs-cl
4 May 2026
Safety

New Mexico child safety trial: New Mexico asks a judge to declare Meta a public nuisance and to order it to pay $3.7B and overhaul its apps to protect children (Diana Novak Jones/Reuters)

DGX agent

Diana Novak Jones / Reuters: New Mexico child safety trial: New Mexico asks a judge to declare Meta a public nuisance and to order it to pay $3.7B and overhaul its apps to protect children — The U.S.

safetytechmeme
4 May 2026
Safety

A profile of BlackBerry's QNX division, whose operating system controls safety features in 275M cars and accounts for half of BlackBerry's revenue (Ben Cohen/Wall Street Journal)

DGX agent

Ben Cohen / Wall Street Journal: A profile of BlackBerry's QNX division, whose operating system controls safety features in 275M cars and accounts for half of BlackBerry's revenue — John Wall has spen

safetytechmeme
3 May 2026
Model Releases

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

DGX agent

arXiv:2604.26959v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into patient-facing healthcare systems offers significant potential to improve access to medical information.

model-releasesarxiv-cs-ai
1 May 2026
Safety

GAVEL: Towards Rule-Based Safety Through Activation Monitoring

DGX agent

arXiv:2601.19768v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be appare

safetyarxiv-cs-ai
1 May 2026
Safety

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely un…

DGX agent

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely unregulated. You just can’t. - Gary The judge just instructed

safetygary-marcus--x
30 Apr 2026
Safety

Test-Time Safety Alignment

DGX agent

arXiv:2604.26167v1 Announce Type: cross Abstract: Recent work has shown that a model's input word embeddings can serve as effective control variables for steering its behavior toward outputs that sati

safetyarxiv-cs-ai
30 Apr 2026
Safety

A Lightweight Explainable Guardrail for Prompt Safety

DGX agent

arXiv:2602.15853v2 Announce Type: replace-cross Abstract: We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly l

safetyarxiv-cs-ai
28 Apr 2026
Safety

Does Machine Unlearning Preserve Clinical Safety? A Risk Analysis for Medical Image Classification

DGX agent

arXiv:2604.23854v1 Announce Type: new Abstract: The application of Deep Learning in medical diagnosis must balance patient safety with compliance with data protection regulations. Machine Unlearning e

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models

DGX agent

arXiv:2511.08484v2 Announce Type: replace Abstract: We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilitie

model-releasesarxiv-cs-ai
28 Apr 2026
Safety

Time-Series Forecasting in Safety-Critical Environments: An EU-AI-Act-Compliant Open-Source Package / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein KI-VO-konformes Open-Source-Paket

DGX agent

arXiv:2604.23859v1 Announce Type: new Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical envir

safetyarxiv-cs-ai
28 Apr 2026
Safety

Australia's eSafety Commissioner issues transparency notices to Roblox, Microsoft's Minecraft, and other online gaming platforms to detail child safety measures (Renju Jose/Reuters)

DGX agent

Renju Jose / Reuters: Australia's eSafety Commissioner issues transparency notices to Roblox, Microsoft's Minecraft, and other online gaming platforms to detail child safety measures — Australia's int

safetytechmeme
22 Apr 2026
Model Releases

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

DGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinct…

DGX agent

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinction risk. We're not. It's not that complicated: AI smarter t

safetyconnor-leahy--x
21 Apr 2026
Safety

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs

DGX agent

arXiv:2509.05367v4 Announce Type: replace-cross Abstract: Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classificati

safetyarxiv-cs-ai
17 Apr 2026
← Previous
1…89101112…297
Next →