💩
A collaborative/competitive speedrun optimizing the fastest algorithm to train a language model to GPT-2-level performance. Achieves 3.28 cross-entropy loss on FineWeb in under 90 seconds and 400M tokens — a 30x speedup over the original llm.c baseline — through novel optimizers, architectures, and systems tricks.