Fetching the paper…
Reading the bibliography…
Deep learning models have become a cornerstone of modern AI research, yet their initializations and learning rates may at times be set in an opaque or ad-hoc fashion due to the high cost of hyperparameter sweeps.
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model, 2019
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E. Dahl, Christopher J. Shallue, and Roger Grosse · 1907
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 1909
Earlier work this paper cites.
Gradient descent: The ultimate optimizer
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, and Erik Meijer · 1909
Earlier work this paper cites.
Greg Yang · 1910
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 1910
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer · 1911
Earlier work this paper cites.
GLU variants improve transformer
Noam Shazeer · 2002
Earlier work this paper cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu · 2002
Earlier work this paper cites.
Tensor Programs II: Neural tangent kernel for any architecture
Greg Yang · 2006
Earlier work this paper cites.
Tensor Programs III: Neural matrix laws
Greg Yang · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences, 2018
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Earlier work this paper cites.
Adaptive input representations for neural language modeling, 2019
Alexei Baevski and Michael Auli · 2019
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training, 2019
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Earlier work this paper cites.
Variable-rate discrete representation learning, 2021
Sander Dieleman, Charlie Nash, Jesse Engel, and Karen Simonyan · 2021
Earlier work this paper cites.
Tensor Programs IV: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Earlier work this paper cites.
Tensor Programs V: Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Cited alongside, same era.
Searching for efficient transformers for language modeling
David So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le · 2021
Cited alongside, same era.
GSPMD: General and scalable parallelization for ML computation graphs
Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman, Yanping Huang, Rahul Joshi, Maxim Krikun, Dmitry Lepikhin, Andy Ly, and Marcello Maggioni et al · 2021
Cited alongside, same era.
RoFormer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, and Quentin Malartic et al · 2023
Later among the works it cites.
Depth dependence of μ \mu p learning rates in relu mlps
Samy Jelassi, Boris Hanin, Ziwei Ji, Sashank J. Reddi, Srinadh Bhojanapalli, and Sanjiv Kumar · 2023
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin, and Cengiz Pehlevan · 2023
Later among the works it cites.
Yiqun Yao and Yequan Wang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zachary Nado, Justin M. Gilmer, Christopher J. Shallue, Rohan Anil, and George E. Dahl · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, and Susannah Young et al · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, and Aidan Clark et al · 2022
Cited alongside, same era.
On the SDEs and scaling rules for adaptive gradient algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora · 2022
Cited alongside, same era.
Self-consistent dynamical field theory of kernel evolution in wide neural networks, 2022
Blake Bordelon and Cengiz Pehlevan · 2022
Cited alongside, same era.
Cramming: Training a language model on a single gpu in one day, 2022
Jonas Geiping and Tom Goldstein · 2022
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Cited alongside, same era.
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Later among the works it cites.
Learning-rate-free learning by d-adaptation, 2023
Aaron Defazio and Konstantin Mishchenko · 2023
Later among the works it cites.
The shaped transformer: Attention models in the infinite depth-and-width limit, 2023
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris Maddison, and Daniel M. Roy · 2023
Later among the works it cites.
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt J. Kusner · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, and Julian Schrittwieser et al · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, and Florian Bressand et al · 2024
Closest in time.
Nemotron-4 15b technical report
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subramanian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, and Ayush Dattagupta et al · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
LLM.int8() and emergent features
Tim Dettmers · 2024
Closest in time.
Grok-1, 2024
XAI · 2024
Closest in time.
Olmo: Accelerating the science of language models, 2024
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, and Yizhong Wang et al · 2024
Closest in time.
Privileged bases in the transformer residual stream
Nelson Elhage, Robert Lasenby, and Christopher Olah · 2024
Closest in time.
Scaling exponents across parameterizations and optimizers
Katie E Everett, Lechao Xiao, Mitchell Wortsman, Alexander A Alemi, Roman Novak, Peter J Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee, and Jeffrey Pennington · 2024
Closest in time.
DeepSeek LLM: Scaling open-source language models with longtermism
DeepSeek-AI, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, and Qiushi Du et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology, 2024
Gemma Team · 2024
Closest in time.
NanoLM: An affordable LLM study benchmark via accurate loss prediction across scales, 2024
Siqi Fan, Xiusheng Huang, Xuezhi Fang, Yiqun Yao, Xiang Li, Ziyi Ni, Xin Jiang, Xuying Meng, Peng Han, and Shuo Shang et al · 2024
Closest in time.
Prodigy: An expeditiously adaptive parameter-free learner, 2024
Konstantin Mishchenko and Aaron Defazio · 2024
Closest in time.
Super consistency of neural network landscapes and learning rate transfer, 2024
Lorenzo Noci, Alexandru Meterez, Thomas Hofmann, and Antonio Orvieto · 2024
Closest in time.
Scaling optimal lr across token horizons, 2024
Johan Bjorck, Alon Benhaim, Vishrav Chaudhary, Furu Wei, and Xia Song · 2024
Closest in time.