Safety

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea

DGX agentx-post
safetydair-ai--x

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New research from Microsoft Research makes the action policy the measured object. 8 models, 6 parallel benchmarks, 41 languages, 2.38M rollouts. Five confounds sit between raw trace similarity and any defensible claim. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, reproducibility caps the gap, and one model asked twice in one language answers differently. Removing all five made the effect larger. Normalised by their own reproducibility, four frontier models each keep 71 to 73% of their action policy across languages. Paper: https://arxiv.org/abs/2608.11110 Track more trending AI papers in our academy: https://academy.dair.ai/

Related

Source: DAIR.AI (X) | 2026-08-12

Loading related sources…