Industry

Distributed training is hard. We adopted DTensor at Runway to prevent silent gradient bugs and it delivered. But we traded performance for c…

Distributed training is hard. We adopted DTensor at Runway to prevent silent gradient bugs and it delivered. But we traded performance for correctness, hitting dispatch overhead, recompilation storms,

DGX agentx-post
industrycristobal-valenzuela--x

Distributed training is hard. We adopted DTensor at Runway to prevent silent gradient bugs and it delivered. But we traded performance for correctness, hitting dispatch overhead, recompilation storms, and MFU drops. Wrote up what we learned and how we work around it. https://runwayml.com/news/dtensor-distributed-training

Source: Cristobal Valenzuela (X) | 2026-05-18

Loading related sources…