Developer builds a 12M-parameter ML framework from scratch in four months
Original titleGoing from knowing zero about ML to building a ML framework from scratch in 4 months is kinda crazy 👀
AISummary
A developer trained a 12M-parameter LLM on a custom ML framework built with a Rust backend and CUDA kernels, including Flash Attention, fused LayerNorm, and fused GELU. The framework claims 3x throughput gains, WebGPU fallback for non-NVIDIA devices, and a TypeScript API installable via npm.
Source: NVIDIA AI Developer · x.comPublished · added here