BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
DGX agentarXiv:2509.02655v3 Announce Type: replace-cross Abstract: Many AI alignment discussions of 'runaway optimisation' focus on RL agents: unbounded utility maximisers that over-optimise a proxy objective