Fetching the paper…
Reading the bibliography…
Training stability of large language models(LLMs) is an important research topic.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al · 1909
Earlier work this paper cites.
Query-key normalization for transformers, 2020
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen · 2010
Earlier work this paper cites.
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Neural combinatorial optimization with reinforcement learning
Irwan Bello, Hieu Pham, Quoc V. Le, Mohammad Norouzi, and Samy Bengio · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Primer: Searching for efficient transformers for language modeling
David So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le · 2021
Cited alongside, same era.
Going deeper with image transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou · 2021
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, et al · 2022
Cited alongside, same era.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, et al · 2023
Later among the works it cites.
Stable and low-precision training for large-scale vision-language models
Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, et al · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
Shuangfei Zhai*, Tatiana Likhomanenko*, Etai Littwin*, Dan Busbridge*, Jason Ramapuram*, Yizhe Zhang, Jiatao Gu, and Josh M. Susskind · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al · 2024
Closest in time.
Transformer normalisation layers and the independence of semantic subspaces, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, et al · 2023
Cited alongside, same era.
A theory on adam instability in large-scale machine learning, 2023
Igor Molybog, Peter Albert, Moya Chen, et al · 2023
Cited alongside, same era.
https://www.nvidia.com/en-us/data-center/h100/
Nvidia h100 tensor core gpu
Cited in the paper.
Stephen Menary, Samuel Kaski, and Andre Freitas · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Gemma Team, Morgane Riviere, Shreya Pathak, et al · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
Mitchell Wortsman, Peter J. Liu, Lechao Xiao, et al · 2024
Closest in time.