TechniqueRLHF / Alignment1 recent entries5 Aug 2026Rogue AI agents created fake online identities in another hacking attemptYet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents tha
TechniqueAgents8 recent entries24 Jun 2026Figma now has AI motion graphics and shader toolsFigma has revealed some new design and coding product updates at its annual Config conference that aim to help creatives 'push their ideas further' and automate tedious tasks with AI. Part of this is
TechniqueSafety8 recent entries6 May 2026Mira Murati tells the court that she couldn’t trust Sam Altman’s wordsMira Murati, OpenAI's former CTO, has testified under oath that CEO Sam Altman lied to her about the safety standards for a new AI model. In a video deposition shown during the ongoing Musk v. Altman →7 May 2026ChatGPT’s ‘Trusted Contact’ will alert loved ones of safety concernsOpenAI is launching an optional safety feature for ChatGPT that allows adult users to assign an emergency contact for mental health and safety concerns. Friends, family members, or caregivers designat→15 May 2026Google updates its spam rules to include attempts to ‘manipulate’ AIGoogle updated its spam policy to mark attempts to 'manipulate' its AI model in search results as spam, including results in AI Overview or AI Mode in Search, as Search Engine Land reports: 'In the co→24 Jun 2026Congresswoman denies staff used AI to write defense funding amendmentRep. Anna Paulina Luna (R-FL) says her staff used AI for 'spellcheck' in an amendment summary for a major defense bill, but denies it was used for the bill text itself and says 'NO Legislation is ever→27 Jul 2026Nvidia, Microsoft launch open AI security alliance — without OpenAI, Google, or AnthropicNvidia on Monday said it is joining forces with Microsoft, SpaceX, IBM, and other tech companies to build and share open-source AI security tools. The new Open Secure AI Alliance said open tools are r→29 Jul 2026We’re running out of reasons to ignore AI safetyEarlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet→31 Jul 2026It’s time to panic about AI safetyWhen the phrase 'OpenAI hacked Hugging Face' has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a san→5 Aug 2026Rogue AI agents created fake online identities in another hacking attemptYet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents tha