Safety
Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4
This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat
This newsletter covers three main topics: advances in automating alignment research to improve AI safety processes, a safety evaluation study of a Chinese AI model, and technical details about HiFloat4, a floating-point format or optimization technique. The issue explores emerging methods to scale alignment research efforts and includes international perspectives on AI safety evaluation practices.
Related
- Anthropic details using AI agents to accelerate alignment research on 'weak-to-strong supervision', where a weak model supervises the training of a stronger one (Anthropic)
- What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
- Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs
- The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
- Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
Source: Import AI | 2026-04-20