The Closing of the Recipe
Part 5. When every model shares one skeleton, the moat becomes the thing nobody publishes.
A six part field guide. Start with the overview or jump straight to any part.
Part 5. When every model shares one skeleton, the moat becomes the thing nobody publishes.
Part 4. Why the most cited recipe in pretraining optimizes the wrong thing, and why emergence is mostly a measurement artifact.
Token-level KV cache eviction can report large logical savings while returning little usable GPU memory to a paged allocator. How much it actually returns is...
A reproducible mechanism study of schedulability, KV-cache lifecycle, and useful tokens per dollar in LLM serving, with a controlled scheduler ablation and h...
Part 3. Why the frontier is gated by memory bandwidth, packaging capacity, and depreciation, and barely at all by ideas about intelligence.
Part 2. Why every frontier model has converged on the same skeleton, and why the silicon, not the science, is what chose it.
Part 1. Why the corpus, not the architecture, is the moat, and why the two loudest panics about running out of data are both told wrong.
A field guide to pretraining for engineers, and why most of what you have read about it is already obsolete.
Thoughts on a possible path for scaling HBM bandwidth through system-level reliability