Back to NewsMini-ML: 12M Parameter LLM Trained From Scratch in RustApril 21, 2026ResearchThey trained a 12M parameter LLM on their own ML framework using a Rust backend and CUDA kernels for flash attention, AdamW, and more. Inspirational project for anyone who wants to better understand how to build LLMs. Sources: tweetRelatedLLM Fused with Mini Computer: Switching Between Text and Machine CodeLLM Architecture GalleryGradMem: Writing Context into LLM Memory via Test-Time Gradient DescentPrevXiaomi Releases MiMo-V2.5Next13+ Attention Mechanisms You Should Know