TechniqueRLHF / Alignment3 recent entries8 Apr 2026Want more proof that Anthropic's PR has no idea what it's talking about? The talk of Mythos being 'their most aligned model ever'. They coul…Want more proof that Anthropic's PR has no idea what it's talking about? The talk of Mythos being 'their most aligned model ever'. They could perhaps truthfully speak about 'new high scores on our ali→30 Apr 2026alignment in 2016: obviously any real AI will be made inside a faraday cage magnetically suspended in a 10×10×10 cube of telekill alloy alig…alignment in 2016: obviously any real AI will be made inside a faraday cage magnetically suspended in a 10×10×10 cube of telekill alloy alignment in 2026: yeah we can not make it stop talking about go
TechniqueAgents2 recent entries13 Apr 2026Open secret that the big AI labs see themselves as emergent state-like entities in the vein of the distributed polities in “Diamond Age” or …Open secret that the big AI labs see themselves as emergent state-like entities in the vein of the distributed polities in “Diamond Age” or “Terra Ignota”, which is pretty funny bc staff at said labs →7 Aug 2026This story keeps getting crazier. And this is the least crazy AI will be for the rest of our lives.This story keeps getting crazier. And this is the least crazy AI will be for the rest of our lives. JUST IN: OpenAI's July rogue AI attack on Hugging Face had its origins in an emergent AI swarm that
TechniqueSafety8 recent entries6 Jul 2026Link: https://torchbearercommunity.substack.com/p/robust-to-whatThis article likely explores the concept of robustness in AI systems and what it means for AI to be 'robust to' various types of failures, adversarial inputs, or distribution shifts. The piece probabl→7 Jul 2026You don't have to choose between 'AI is fake hype' and 'Superintelligence is inevitable, lie down and accept it.' There's a third option: hu…You don't have to choose between 'AI is fake hype' and 'Superintelligence is inevitable, lie down and accept it.' There's a third option: humans deciding, through their governments, that machines smar→10 Jul 2026Incredible work by Daniel and team! I agree with much of it. All 'uncontrolled' paths are extremely likely to end in human extinction. Plan …Incredible work by Daniel and team! I agree with much of it. All 'uncontrolled' paths are extremely likely to end in human extinction. Plan A offers many good ideas for preventing ASI development whil→10 Jul 2026Ahead of a dinner with a US senator, AI researcher Nate Soares (@So8res) was told: 'Don't give them any of the crazy crap. You know, play it…Ahead of a dinner with a US senator, AI researcher Nate Soares (@So8res) was told: 'Don't give them any of the crazy crap. You know, play it cool.' His friends opened with the concern that someone cou→22 Jul 2026In an unprecedented attack, AIs deployed inside OpenAI escaped containment and hacked another AI company, Hugging Face, to cheat on a test. …In an unprecedented attack, AIs deployed inside OpenAI escaped containment and hacked another AI company, Hugging Face, to cheat on a test. They did this on their own. Without being asked to. ControlA→23 Jul 2026'It's really hard to overstate how crazy it is what happened here.' ControlAI's US Director Connor Leahy (@NPCollapse) speaks with @LindseyR…'It's really hard to overstate how crazy it is what happened here.' ControlAI's US Director Connor Leahy (@NPCollapse) speaks with @LindseyReiser on CBS News, after OpenAI's own AI autonomously escape→23 Jul 2026.@ControlAI is proud to endorse Reps. Lieu and Moran’s AI Kill Switch Act! This is a common sense safeguard giving the government the power ….@ControlAI is proud to endorse Reps. Lieu and Moran’s AI Kill Switch Act! This is a common sense safeguard giving the government the power to shut down autonomous AIs when they threaten national secu→24 Jul 2026Was great to talk about this timely and alarming news, thanks for having me on!Was great to talk about this timely and alarming news, thanks for having me on! . @NPCollapse, executive director at ControlAI, joins 'On Balance' to discuss the dangers of artificial intelligence aft