So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.
Zyphra's correlated expert placement cuts MoE token copies by 58%
AISummary
Zyphra places experts that are frequently chosen together on the same GPU, so each token is sent to a GPU only once regardless of how many of its experts reside there. On 8 GPUs, this removes up to 58% of token copies sent across the network.
Post on XView on X
Source: Zyphra · x.comPublished · added here
