Safety
New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I hig…
New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I highly recommend giving it a read. Link to the paper: https://a
New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I highly recommend giving it a read. Link to the paper: https://arxiv.org/pdf/2603.16928 Can LLM agents figure out that their "private" chain of thought is being monitored, with no hints, no training, just from getting blocked? Yes. And some then explicitly reason to hide their reasoning. New paper with @FlemmingKondrup (co-first authors) & @Yoshua_Bengio: "Noticing …
Related
- Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
- Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
Source: Yoshua Bengio (X) | 2026-07-14