Skip to content
Trending storyDeveloping

Zyphra claims 1.16x–2.63x faster MoE token exchange in Megatron-LM, unverified

2 articles1 sourcesince Oct 8Last article Yesterday ·

Overview

AISummary of one article

Zyphra reports that its MoE training optimizations speed up token exchange by 1.16x to 2.63x and full training steps by up to 1.41x in Megatron-LM on 8 to 64 GPUs.

The gains are largest when each token uses more experts and those experts span several nodes.

Written by AI from one article, by Zyphra

Check the sources:

Developments

2 developments

  1. Oct 8, 11:35 AM ET · 1 article
    Zyphra describes a lossless approach that reduces communication in MoE training by reorganizing expert and token placement
    Zyphra: Zyphra's lossless method cuts communication for MoE expert routing
  2. Oct 8, 11:35 AM ET · 1 article
    Zyphra reports MoE token exchange speedups of 1.16x to 2.63x in Megatron-LM
    Zyphra: Zyphra reports up to 2.63x faster MoE token exchange in Megatron-LM

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Zyphra
    Zyphra's lossless method cuts communication for MoE expert routing

    AIZyphra says its approach is lossless: the same tokens still reach the same experts, with unchanged architecture, routing decisions, and training objective. By reorganizing where experts and tokens live, it reduces the communication needed to perform the same computation.

  2. Zyphra
    Zyphra reports up to 2.63x faster MoE token exchange in Megatron-LM

    AIZyphra reports that its MoE training optimizations speed up token exchange by 1.16x to 2.63x and full training steps by up to 1.41x in Megatron-LM on 8 to 64 GPUs. The gains are largest when each token uses more experts and those experts span several nodes.

Heat trend

Not enough continuous observations to show a trend yet.