Single-GPU mixture-of-experts LLM trained from scratch in 8 days
Original titleNice showcase that interesting LLM work can be done on single GPU!
AISummary
Giles Thomas extended the GPT-2-style code from Sebastian Raschka's "Build a Large Language Model (from Scratch)" into a 6-expert, 2-active mixture-of-experts model and trained it from scratch over 8 days. Raschka praised the project as interesting LLM work done on a single GPU.
Source: Sebastian Raschka · x.comPublished · added here