AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
49+ results
Safety

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? …

DGX agent

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? This link is from Nov 2023 & shows how tech CEOs used 'safet

safetygary-marcus--x
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Worth remembering when @janleike quit over OpenAI safety concerns.

DGX agent

Jan Leike, OpenAI's safety and alignment lead, resigned in May 2024, citing concerns about the company's commitment to safety practices and the deprioritization of safety work relative to product deve

safetygary-marcus--x
7 May 2026
Safety

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score…

DGX agent

Tesla Insurance update With the latest version of Safety Score (v3.0), every mile you drive with FSD Supervised enabled will receive a score of 100. This allows you to maintain a higher average safety

safetyelon-musk--x
14 Apr 2026
Safety

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot…

DGX agent

Sam Altman: “Concentration of power with AI is a terrifying thing.” 'A lot of the talk about safety concerns is well-founded, and then a lot of it is about people that just really, even if it's slight

safetyclem-delangue--x
28 Jul 2026
Safety

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire

DGX agent

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire Altman’s AI safety proposal: let us win, or everybody loses https://ft.trib.al/UrDI

safetygary-marcus--x
2 Jul 2026
Safety

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent…

DGX agent

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false

safetyclem-delangue--x
21 Apr 2026
Safety

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.…

DGX agent

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running

safetyopenai--x
20 Jul 2026
Safety

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

DGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

safetydair-ai--x
1 May 2026
Safety

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance t…

DGX agent

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance to adoption. In this op-ed in the @FT, I underscore that Euro

safetyyoshua-bengio--x
27 Apr 2026
Safety

I have found that asking for a sestina regularly triggers Opus 4.7's safety guardrails. The forbidden poetic form!

DGX agent

Ethan Mollick reported that requesting Claude Opus 4.7 to write sestinas—a complex poetic form with strict structural requirements—frequently triggers the model's safety guardrails, suggesting the AI

safetyethan-mollick--x
16 Apr 2026
Safety

it’s just for your safety

DGX agent

On July 23, 2026, a Polymarket tweet announced that Anthropic donated $20 million to a political nonprofit calling for stricter AI regulation ahead of the U.S. midterm elections. The post, posted at 4

safetyyann-lecun--x
23 Jul 2026
Safety

Major improvement in safety statistics with FSD turned on in the Netherlands!

DGX agent

Elon Musk posted on X claiming that Tesla's Full Self-Driving (FSD) system resulted in significant improvements to vehicle safety statistics in the Netherlands. The post suggests FSD activation correl

safetyelon-musk--x
9 Jun 2026
Safety

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

DGX agent

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

safetyethan-mollick--x
5 Aug 2026
Safety

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on a…

DGX agent

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on as an afterthought. Because trust is everything when robots s

safetyyann-lecun--x
7 Jul 2026
Safety

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now s…

DGX agent

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now so it’s taking things slow and smoothly. Highway performance

safetyelon-musk--x
29 Jun 2026
Safety

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know wh…

DGX agent

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know what risks everyone will face if/when open source reaches Myth

safetyethan-mollick--x
25 Jun 2026
Safety

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alig…

DGX agent

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alignment. You can read it here: https://arxiv.org/abs/2606.1691

safetyyoshua-bengio--x
22 Jun 2026
Safety

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a…

DGX agent

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a different kind of answer about what is real and what is mark

safetygary-marcus--x
10 Jun 2026
Safety

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. …

DGX agent

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't d

safetyyann-lecun--x
10 Jun 2026
Safety

fable’s safety guardrails broken within an hour. raise your hand if you are surprised.

DGX agent

fable’s safety guardrails broken within an hour. raise your hand if you are surprised. We tested Anthropic’s new @claudeai Fable 5. It did not fail like an ordinary jailbreak. It failed more quietly.

safetygary-marcus--x
9 Jun 2026
Safety

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Ye…

DGX agent

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Yet most of the field kept licking the boot of the companies a

safetyconnor-leahy--x
5 Jun 2026
Safety

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the o…

DGX agent

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the open on @huggingface, so researchers everywhere can scrutiniz

safetyclem-delangue--x
30 May 2026
Safety

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_!

DGX agent

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_! Is AI development progressing too quickly? @business' @shiringh

safetyyoshua-bengio--x
26 May 2026
Safety

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be ab…

DGX agent

I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea

safetygary-marcus--x
23 May 2026
Safety

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks t…

DGX agent

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks things are fine readiness wise or incentive wise etc. https:/

safetygary-marcus--x
22 May 2026
Safety

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can fi…

DGX agent

We also had three third-party AI safety organizations provide feedback on our analysis: @redwood_ai, @apolloaievals, @METR_Evals. You can find @redwood_ai's report here: https://blog.redwoodresearch.o

safetyopenai--x
8 May 2026
Safety

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founde…

DGX agent

Many, many OpenAI employees quit over safety concerns, including @DKokotajlo, William Saunders, @sjgadler, etc as well @Janleike. The founders of Anthropic such as @DarioAmodei and @jackclarkSF may ha

safetygary-marcus--x
7 May 2026
Safety

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely un…

DGX agent

Dear @elonmusk, If you still genuinely care about AI safety, you can’t let the Trump administration leave the AI industry almost entirely unregulated. You just can’t. - Gary The judge just instructed

safetygary-marcus--x
30 Apr 2026
Safety

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinct…

DGX agent

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinction risk. We're not. It's not that complicated: AI smarter t

safetyconnor-leahy--x
21 Apr 2026
Safety

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this pow…

DGX agent

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this powerful technology, not provide a get-out-of-jail-free card ag

safetygary-marcus--x
14 Apr 2026
Safety

Given the extremely high rate of recidivism, this is important for community safety

DGX agent

Given the extremely high rate of recidivism, this is important for community safety Murder registry is a good idea from Elon. People should know if they are close to a murderer so proper precautions c

safetyelon-musk--x
11 Apr 2026
Safety

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupe…

DGX agent

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupervised and complex situations. 600 miles in with FSD v14.3 a

safetyelon-musk--x
9 Apr 2026
Model Releases

🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 http://mistral.ai/news/shie…

DGX agent

Mistral AI introduced Shieldstral, a 3‑billion‑parameter, open‑weights model designed for content‑safety tasks and capable of on‑device deployment. The announcement was shared via a tweet from @Mistra

model-releasesmistral-ai--x
4 Aug 2026
Model Releases

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

DGX agent

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

model-releasesdair-ai--x
12 Jul 2026
Model Releases

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety resea…

DGX agent

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety research. So much more to do. We are 1% done. We've put together

model-releasesboris-cherny--x
6 Jul 2026
Model Releases

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repea…

DGX agent

GPT‑5.6 Sol launches with our most robust safety stack yet. We strengthened real-time protections against high-risk cyber activity and repeated misuse, then spent weeks hardening the system with human

model-releasesopenai--x
26 Jun 2026
Model Releases

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind.

DGX agent

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind. I found another API that offers claude-oceanus-v1-p the pricing and t

model-releasesclem-delangue--x
4 Jun 2026
Model Releases

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍

DGX agent

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍 JUST IN: Anthropic announces it will roll out Claude Mythos “in the com

model-releasesjeremy-howard--x
28 May 2026
Model Releases

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a…

DGX agent

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a human would never try on another human. In one run, GPT-5.1

model-releasesemad-mostaque--x
13 May 2026
Model Releases

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

DGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

model-releasesjerry-liu--x
11 May 2026
Safety

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to se…

DGX agent

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to see a debunking of their data, then we safety advocates should

safetyethan-mollick--x
12 Apr 2026
Safety

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

DGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

safetyopenai--x
15 Jul 2026
Safety

but they won’t.

DGX agent

but they won’t. AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better

safetygary-marcus--x
25 Jul 2026
Safety

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, man…

DGX agent

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, many forefront members of the AI-Safety community, in their fer

safetyyann-lecun--x
30 Jul 2026
Safety

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and…

DGX agent

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety an

safetyclem-delangue--x
24 Jul 2026
Safety

Q1 2026 Shareholder Update https://ir.tesla.com/#quarterly-disclosure We continued to make meaningful progress on the build out of the infra…

DGX agent

Q1 2026 Shareholder Update https://ir.tesla.com/#quarterly-disclosure We continued to make meaningful progress on the build out of the infrastructure & AI software that underpins our Robotaxi & future

safetyelon-musk--x
22 Apr 2026
Safety

4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctne…

DGX agent

4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctness, regression safety, and minimality. Basically, get the LL

safetyai21-labs--x
15 Apr 2026
Safety

This Meta campaign is a case study in strategic reframing, with 4 major examples: 1) REFRAMES THE AI RACE FROM “who builds it the best” TO “…

DGX agent

This Meta campaign is a case study in strategic reframing, with 4 major examples: 1) REFRAMES THE AI RACE FROM “who builds it the best” TO “who distributes it to the most people”: Instead of fighting

safetyyann-lecun--x
10 Aug 2026
← Previous
1
Next →
1,620 results
← Previous
123…34
Next →