AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
61+ results
21 Apr 2026

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? …

SafetyDGX agent

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? This link is from Nov 2023 & shows how tech CEOs used 'safet

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent…

SafetyDGX agent

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinct…

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinction risk. We're not. It's not that complicated: AI smarter t

7 May 2026

Worth remembering when @janleike quit over OpenAI safety concerns.

SafetyDGX agent

Jan Leike, OpenAI's safety and alignment lead, resigned in May 2024, citing concerns about the company's commitment to safety practices and the deprioritization of safety work relative to product deve

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founde…

SafetyDGX agent

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founders of Anthropic such as @DarioAmodei and @jackclarkSF may ha

14 Apr 2026

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score…

SafetyDGX agent

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score of 100. This allows you to maintain a higher average safety

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this pow…

SafetyDGX agent

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this powerful technology, not provide a get-out-of-jail-free card ag

28 Jul 2026

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot…

SafetyDGX agent

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot of it is about people that just really, even if it's slight

2 Jul 2026

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire

SafetyDGX agent

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire Altman’s AI safety proposal: let us win, or everybody loses https://ft.trib.al/UrDI

20 Jul 2026

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.…

SafetyDGX agent

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running

1 May 2026

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

SafetyDGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

27 Apr 2026

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance t…

SafetyDGX agent

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance to adoption. In this op-ed in the @FT, I underscore that Euro

If extremely violent criminals are not imprisoned, eventually they will murder innocent people

SafetyDGX agent

If extremely violent criminals are not imprisoned, eventually they will murder innocent people This bodega owner told ABC a year ago that he fears for his safety in NY Last night, he was kiIIed by a s

16 Apr 2026

I have found that asking for a sestina regularly triggers Opus 4.7's safety guardrails. The forbidden poetic form!

SafetyDGX agent

Ethan Mollick reported that requesting Claude Opus 4.7 to write sestinas—a complex poetic form with strict structural requirements—frequently triggers the model's safety guardrails, suggesting the AI

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Te…

SafetyDGX agent

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Tesla-driven robotaxis. • Waymo: 58 incidents • Zoox: 3 incide

23 Jul 2026

it’s just for your safety

SafetyDGX agent

On July 23, 2026, a Polymarket tweet announced that Anthropic donated $20 million to a political nonprofit calling for stricter AI regulation ahead of the U.S. midterm elections. The post, posted at 4

9 Jun 2026

Major improvement in safety statistics with FSD turned on in the Netherlands!

SafetyDGX agent

Elon Musk posted on X claiming that Tesla's Full Self-Driving (FSD) system resulted in significant improvements to vehicle safety statistics in the Netherlands. The post suggests FSD activation correl

fable’s safety guardrails broken within an hour. raise your hand if you are surprised.

SafetyDGX agent

fable’s safety guardrails broken within an hour. raise your hand if you are surprised. We tested Anthropic’s new @claudeai Fable 5. It did not fail like an ordinary jailbreak. It failed more quietly.

5 Aug 2026

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

SafetyDGX agent

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-to…

SafetyDGX agent

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

SafetyDGX agent

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

7 Jul 2026

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on a…

SafetyDGX agent

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on as an afterthought. Because trust is everything when robots s

29 Jun 2026

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now s…

SafetyDGX agent

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now so it’s taking things slow and smoothly. Highway performance

25 Jun 2026

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know wh…

SafetyDGX agent

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know what risks everyone will face if/when open source reaches Myth

22 Jun 2026

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alig…

SafetyDGX agent

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alignment. You can read it here: https://arxiv.org/abs/2606.1691

10 Jun 2026

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a…

SafetyDGX agent

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a different kind of answer about what is real and what is mark

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. …

SafetyDGX agent

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't d

5 Jun 2026

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Ye…

SafetyDGX agent

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Yet most of the field kept licking the boot of the companies a

30 May 2026

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the o…

SafetyDGX agent

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the open on @huggingface, so researchers everywhere can scrutiniz

26 May 2026

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_!

SafetyDGX agent

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_! Is AI development progressing too quickly? @business' @shiringh

23 May 2026

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

SafetyDGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

22 May 2026

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks t…

SafetyDGX agent

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks things are fine readiness wise or incentive wise etc. https:/

8 May 2026

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can fi…

SafetyDGX agent

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can find @redwood_ai's report here: https://blog.redwoodresearch.o

Good summary of today, @katiemiller, but then again it is getting hard to track of the total number of ex-board members who have called Altm…

SafetyDGX agent

Good summary of today, @katiemiller, but then again it is getting hard to track of the total number of ex-board members who have called Altman a liar 🤷‍♂️ Also hard to keep track how many OpenAI safet

30 Apr 2026

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely un…

SafetyDGX agent

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely unregulated. You just can’t. - Gary The judge just instructed

11 Apr 2026

Given the extremely high rate of recidivism, this is important for community safety

SafetyDGX agent

Given the extremely high rate of recidivism, this is important for community safety Murder registry is a good idea from Elon. People should know if they are close to a murderer so proper precautions c

9 Apr 2026

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupe…

SafetyDGX agent

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupervised and complex situations. 600 miles in with FSD v14.3 a

4 Aug 2026

🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 http://mistral.ai/news/shie…

Model ReleasesDGX agent

Mistral AI introduced Shieldstral, a 3‑billion‑parameter, open‑weights model designed for content‑safety tasks and capable of on‑device deployment. The announcement was shared via a tweet from @Mistra

The model takes moderation policy as a plain-language question and returns a calibrated score. Text and images — one interface. Read the ful…

SafetyDGX agent

Mistral AI has released Shieldstral, a 3‑billion‑parameter open‑weight model for content safety that can run locally on device. It interprets plain‑language moderation queries and returns calibrated s

Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles. Beyond seeing, Alpamayo understands and …

SafetyDGX agent

Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles. Beyond seeing, Alpamayo understands and reasons through the complex world - thinks before it acts. I

12 Jul 2026

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

Model ReleasesDGX agent

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

6 Jul 2026

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety resea…

Model ReleasesDGX agent

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety research. So much more to do. We are 1% done. We've put together

26 Jun 2026

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repea…

Model ReleasesDGX agent

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human

4 Jun 2026

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind.

Model ReleasesDGX agent

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind. I found another API that offers claude-oceanus-v1-p the pricing and t

28 May 2026

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍

Model ReleasesDGX agent

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍 JUST IN: Anthropic announces it will roll out Claude Mythos “in the com

13 May 2026

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a…

Model ReleasesDGX agent

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a human would never try on another human. In one run, GPT-5.1

11 May 2026

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

Model ReleasesDGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

12 Apr 2026

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to se…

SafetyDGX agent

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to see a debunking of their data, then we safety advocates should

15 Jul 2026

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

SafetyDGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

25 Jul 2026

but they won’t.

SafetyDGX agent

but they won’t. AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better

30 Jul 2026

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, man…

SafetyDGX agent

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, many forefront members of the AI-Safety community, in their fer

Cohere has joined @NVIDIA alongside industry leaders in founding the Open Secure AI Alliance. Everybody should have the capability to keep t…

SafetyDGX agent

Cohere has joined @NVIDIA alongside industry leaders in founding the Open Secure AI Alliance. Everybody should have the capability to keep their infrastructure secure. Everybody deserves access to mod

24 Jul 2026

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and…

SafetyDGX agent

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety an

22 Apr 2026

Q1 2026 Shareholder Update https://ir.tesla.com/#quarterly-disclosure We continued to make meaningful progress on the build out of the infra…

SafetyDGX agent

Q1 2026 Shareholder Update https://ir.tesla.com/#quarterly-disclosure We continued to make meaningful progress on the build out of the infrastructure & AI software that underpins our Robotaxi & future

15 Apr 2026

4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctne…

SafetyDGX agent

4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctness, regression safety, and minimality. Basically, get the LL

10 Aug 2026

This Meta campaign is a case study in strategic reframing, with 4 major examples: 1) REFRAMES THE AI RACE FROM “who builds it the best” TO “…

SafetyDGX agent

This Meta campaign is a case study in strategic reframing, with 4 major examples: 1) REFRAMES THE AI RACE FROM “who builds it the best” TO “who distributes it to the most people”: Instead of fighting

12 May 2026

Sam Altman swearing to tell the whole truth, and then failing to do so. May 2023.

SafetyDGX agent

Sam Altman made statements under oath in May 2023 regarding AI safety and OpenAI's practices, but Gary Marcus critiqued these statements as incomplete or misleading, suggesting Altman failed to fully

24 Apr 2026

Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the d…

SafetyDGX agent

Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the delivery of information and entertainment and the frictionles

8 Apr 2026

Chilling. The only thing I got wrong here in @politico was the year. This is exactly where we are now.

SafetyDGX agent

* **What it likely covers:** This entry likely discusses contemporary concerns regarding the safety and restriction of freedom of speech or civil liberties, drawing a parallel between past discus...

22 Jul 2026

It's essential that defenders have the same capabilities as attackers. This is a preview of the future where our ham-fisted safeguards and t…

SafetyDGX agent

It's essential that defenders have the same capabilities as attackers. This is a preview of the future where our ham-fisted safeguards and the doomsday and safety drumbeat make us decidely less safe.

← Previous
1
Next →
1,620 results
← Previous
123…27
Next →