Safety
// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…
// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea
// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New research from Microsoft Research makes the action policy the measured object. 8 models, 6 parallel benchmarks, 41 languages, 2.38M rollouts. Five confounds sit between raw trace similarity and any defensible claim. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, reproducibility caps the gap, and one model asked twice in one language answers differently. Removing all five made the effect larger. Normalised by their own reproducibility, four frontier models each keep 71 to 73% of their action policy across languages. Paper: https://arxiv.org/abs/2608.11110 Track more trending AI papers in our academy: https://academy.dair.ai/
Related
- Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
- // The agent is its own best speculator // Agents spend a large share of wall-clock time waiting on tool results. Speculation hides that lat…
- // Latent Agents // Multi-agent debate makes models reason better. It also burns tokens generating long transcripts before any answer comes …
Source: DAIR.AI (X) | 2026-08-12