Fetching the paper…
Reading the bibliography…
Recently, the Muon optimizer based on matrix orthogonalization has demonstrated strong results in training small-scale language models, but the scalability to larger models has not been proven.
“Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”, 2020
Mohammad Shoeybi et al · 1909
Earlier work this paper cites.
“Fast Transformer Decoding: One Write-Head is All You Need”, 2019
Noam Shazeer · 1911
Earlier work this paper cites.
“Singular value decomposition for genome-wide expression data processing and modeling”
Orly Alter, Patrick. Brown and David Botstein · 2000
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”, 2020
Jared Kaplan et al · 2001
Earlier work this paper cites.
“The effective rank: A measure of effective dimensionality”
Olivier Roy and Martin Vetterli · 2007
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”, 2021
Dan Hendrycks et al · 2009
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension”, 2017
Mandar Joshi et al · 2017
Earlier work this paper cites.
“Preconditioned Stochastic Gradient Descent”
Xi-Lin Li · 2017
Earlier work this paper cites.
“Preconditioner on Matrix Lie Group for SGD”, 2018
Xi-Lin Li · 2018
Earlier work this paper cites.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
“ZeRO: Memory optimizations Toward Training Trillion Parameter Models”
Samyam Rajbhandari et al · 2020
Earlier work this paper cites.
“Program Synthesis with Large Language Models”, 2021
Jacob Austin et al · 2021
Earlier work this paper cites.
“Evaluating Large Language Models Trained on Code”, 2021
Mark Chen et al · 2021
Earlier work this paper cites.
“Training Verifiers to Solve Math Word Problems”, 2021
Karl Cobbe et al · 2021
Earlier work this paper cites.
“Measuring Mathematical Problem Solving With the MATH Dataset”, 2021
Dan Hendrycks et al · 2021
Cited alongside, same era.
“Training Compute-Optimal Large Language Models”, 2022
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“Black Box Lie Group Preconditioners for SGD”, 2022
Xilin Li · 2022
Cited alongside, same era.
“Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them”, 2022
Mirac Suzgun et al · 2022
Cited alongside, same era.
“C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models”, 2023
“CMMLU: Measuring massive multitask language understanding in Chinese”, 2024
Haonan Li et al · 2024
Later among the works it cites.
“Stochastic Hessian Fittings with Lie Groups”, 2024
Xi-Lin Li · 2024
Later among the works it cites.
“Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training”
Hong Liu et al · 2024
Later among the works it cites.
Team OLMo et al · 2024
Later among the works it cites.
“GPT-4 Technical Report”, 2024
OpenAI et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuzhen Huang et al · 2023
Cited alongside, same era.
“CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?”, 2023
Tianwen Wei et al · 2023
Cited alongside, same era.
“Old Optimizer, New Norm: An Anthology”, 2024
Jeremy Bernstein and Laker Newhouse · 2024
Cited alongside, same era.
“Deepseek llm: Scaling open-source language models with longtermism”
Xiao Bi et al · 2024
Cited alongside, same era.
“Deep Learning Optimizers as Steepest Descent in Normed Spaces”, 2024
Franz Cesista · 2024
Cited alongside, same era.
“DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model”, 2024
DeepSeek-AI · 2024
Cited alongside, same era.
“DeepSeek-V3 Technical Report”, 2024
DeepSeek-AI et al · 2024
Cited alongside, same era.
“The Case for Muon”, 2024
Louis Franz · 2024
Cited alongside, same era.
“Curvature-Informed SGD via General Purpose Lie-Group Preconditioners”, 2024
Omead Pooladzandi and Xi-Lin Li · 2024
Later among the works it cites.
“Gemini: A Family of Highly Capable Multimodal Models”, 2024
Gemini Team et al · 2024
Later among the works it cites.
“Gemma 2: Improving open language models at a practical size”
Gemma Team et al · 2024
Later among the works it cites.
“MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark”, 2024
Yubo Wang et al · 2024
Later among the works it cites.
An Yang et al · 2024
Later among the works it cites.
“MARS: Unleashing the Power of Variance Reduction for Training Large Models”, 2024
Huizhuo Yuan et al · 2024
Later among the works it cites.
“Training Deep Learning Models with Norm-Constrained LMOs”, 2025
Thomas Pethick et al · 2025
Closest in time.
“Kimi k1.5: Scaling Reinforcement Learning with LLMs”, 2025
Kimi Team · 2025
Closest in time.
“SOAP: Improving and Stabilizing Shampoo using Adam”
Nikhil Vyas et al · 2025
Closest in time.
“Jiacheng You’s discussion on Muon’s Update RMS”, 2025
Jiacheng You · 2025
Closest in time.