MoE training spends 45-60% of step time on cross-node token exchange
Original titleMoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and token...
AISummary
In Zyphra's runs, MoE token exchange between experts consumed 13-24% of step time on one node and 45-60% across four nodes. Because experts are spread across GPUs and nodes, tokens must be sent to their experts and returned, making this communication a major training cost as models scale.
Source: Zyphra · x.comPublished · added here