Mayank put in a crazy amount of work to get pretraining to work on 3 gens of Nvidia GPUs and 2 gens of TPUs! Very good model for such a small size
Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs
AISummary
Mayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.
Post on XView on X
@tri_dao
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs. No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase. Meet Rigel 🧵
Source: Tri Dao · x.comPublished · added here
