The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster
The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes...
AISummary
The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster
Source: Zyphra · x.com