Safety

New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I hig…

New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I highly recommend giving it a read. Link to the paper: https://a

DGX agentx-post
safetyyoshua-bengio--x

New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I highly recommend giving it a read. Link to the paper: https://arxiv.org/pdf/2603.16928 Can LLM agents figure out that their "private" chain of thought is being monitored, with no hints, no training, just from getting blocked? Yes. And some then explicitly reason to hide their reasoning. New paper with @FlemmingKondrup (co-first authors) & @Yoshua_Bengio: "Noticing …

Related

Source: Yoshua Bengio (X) | 2026-07-14

Loading related sources…