Skip to content
Trending storyDeveloping

Zyphra says routing-pattern expert placement speeds MoE token communication up to 2.63x on AMD MI300X

3 articles1 sourcesince Oct 8Last article Yesterday ·

Overview

AISummary of one article

Zyphra researchers report that exploiting patterns in how tokens are routed to experts makes Mixture of Experts communication up to 2.63x faster on AMD MI300X GPUs.

The speedup comes without changing the model itself.

Written by AI from one article, by Zyphra

Check the sources:

Developments

3 developments

  1. Oct 8, 11:35 AM ET · 1 article
    Zyphra proposes correlated expert placement to cut MoE token copies across GPUs
    Zyphra: Zyphra's correlated expert placement cuts MoE token copies by 58%
  2. Oct 8, 11:35 AM ET · 1 article
    Zyphra reports expert-pair routing patterns in MoE layers and a method to exploit them for expert placement
    Zyphra: Zyphra finds MoE expert routing has strong, exploitable pair patterns
  3. Oct 8, 11:35 AM ET · 1 article
    Zyphra uses token routing patterns to speed MoE expert communication up to 2.63x on AMD MI300X
    Zyphra: Zyphra speeds MoE token routing up to 2.63x on AMD MI300X GPUs

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Zyphra
    Zyphra's correlated expert placement cuts MoE token copies by 58%

    AIZyphra places experts that are frequently chosen together on the same GPU, so each token is sent to a GPU only once regardless of how many of its experts reside there. On 8 GPUs, this removes up to 58% of token copies sent across the network.

  2. Zyphra
    Zyphra finds MoE expert routing has strong, exploitable pair patterns

    AIIn one layer, just 64 of 8,128 possible expert pairs handled 42% of tokens, far above the 1.6% expected if experts were chosen independently. Zyphra says this structure can be exploited to improve expert placement.

  3. Zyphra
    Zyphra speeds MoE token routing up to 2.63x on AMD MI300X GPUs

    AIZyphra researchers report that exploiting patterns in how tokens are routed to experts makes Mixture of Experts communication up to 2.63x faster on AMD MI300X GPUs. The speedup comes without changing the model itself.

Heat trend

Not enough continuous observations to show a trend yet.