Skip to content
View original post on X: Tri Dao· 44/100AI score44/100

Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

AISummary

Mayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

Post on XView on X
@tri_dao

Mayank put in a crazy amount of work to get pretraining to work on 3 gens of Nvidia GPUs and 2 gens of TPUs! Very good model for such a small size

Mayank Mishra@MayankMish98
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs. No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase. Meet Rigel 🧵
View quoted post on X

Source: Tri Dao · x.comPublished · added here