We do not believe AdamW or Muon is the most efficient update rule that could exist for language-model training. The space of possible algorithms is much larger than the set in common use. Our thesis is that accurate measurement and sustained incentives can make exploring that space productive.
Some foundational methods remain central across generations of models. That persistence motivates us to revisit them. Research such as FlashAttention shows how rethinking a low-level operation can improve the efficiency of an existing architecture; work scaling Muon to LLM training shows that optimizer development remains an active source of progress.
Academic groups, open-source communities, independent researchers, and frontier labs all contribute to this work. Refinery asks what happens when a decentralized network has a clear objective, a credible test, and a direct reason to keep searching.
Research context: FlashAttention · Muon is Scalable for LLM Training