In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime.
Original titleIn Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime.
AISummary
At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.
Source: Zyphra · x.com