Industry
Distributed training is hard. We adopted DTensor at Runway to prevent silent gradient bugs and it delivered. But we traded performance for c…
Distributed training is hard. We adopted DTensor at Runway to prevent silent gradient bugs and it delivered. But we traded performance for correctness, hitting dispatch overhead, recompilation storms,
Distributed training is hard. We adopted DTensor at Runway to prevent silent gradient bugs and it delivered. But we traded performance for correctness, hitting dispatch overhead, recompilation storms, and MFU drops. Wrote up what we learned and how we work around it. https://runwayml.com/news/dtensor-distributed-training
Source: Cristobal Valenzuela (X) | 2026-05-18