TechniqueRLHF / Alignment3 recent entries20 Apr 2026Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat→24 Apr 2026Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build …Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of the
TechniqueAgents7 recent entries13 Apr 2026Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowermentImport AI issue 453 covers research and developments around vulnerabilities in AI agent systems, including methods for breaking or adversarially manipulating AI agents. The issue also features MirrorC→13 Apr 2026As AI agents accelerate coding, what is the future of software engineering? Some trends are clear, such as the Product Management Bottleneck…As AI agents accelerate coding, what is the future of software engineering? Some trends are clear, such as the Product Management Bottleneck, referring to the idea that we are more constrained by deci→13 Apr 2026A new standard for research: How UC Riverside is securing the path to federal grants with Google Public SectorAt the University of California, Riverside (UCR), scientific breakthroughs depend on quickly moving from a hypothesis to a finished study. Yet for many researchers, the path to federal grants is often→15 Apr 2026The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to a…The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to an agent browsing a similar index in vastly different ways. A→17 Apr 2026I went on @BBCNewsnight this week to discuss the recent developments in AI's capabilities, as well as the potential harms and concentration …I went on @BBCNewsnight this week to discuss the recent developments in AI's capabilities, as well as the potential harms and concentration of power they could entail. We need coordinated internationa→20 Apr 2026Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat→24 Apr 2026Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build …Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of the
TechniqueFine-tuning1 recent entries20 Apr 2026Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat
TechniqueSafety8 recent entries24 Apr 2026Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build …Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of the→4 May 2026Import AI 455: AI systems are about to start building themselves.This newsletter entry discusses advancements in automating AI research and development processes, exploring how artificial intelligence systems are becoming capable of autonomously designing and impro→11 May 2026Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computerThis newsletter covers three main topics: the relationship between AI capabilities relative to human intelligence (RSI) and its potential economic impacts, regulatory approaches that emphasize flexibi→18 May 2026Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignmentThis newsletter explores three AI topics: potential weaponization risks of AI systems (framed as 'AI stuxnet'), issues with the Muon optimizer in machine learning, and progress or approaches toward po→26 May 2026Import AI 458: Reckoning with the future; and a singularity storyImport AI 458 discusses perspectives on AI's future trajectory and potential long-term scenarios, likely including analysis of singularity concepts and their implications. The newsletter entry examine→1 Jun 2026Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systemsThis newsletter issue discusses three key topics in AI development and safety: the challenges involved in overseeing and controlling advanced AI systems, empirical findings about how protein folding A→8 Jun 2026Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racingThis Import AI newsletter issue covers three main topics: societal implications of reward hacking (optimizing for measurable metrics at the expense of intended goals), new reinforcement learning data →22 Jun 2026Import AI 462: Superpersuasion; self-sustaining AI; paths to ASIThis newsletter issue covers three major AI topics: techniques for making AI systems more persuasive ('superpersuasion'), the concept of self-sustaining or self-improving AI systems, and various theor