Fetching the paper…
Reading the bibliography…
We propose a novel neural network architecture, the normalized Transformer (nGPT) with representation learning on the hypersphere.
Animating rotation with quaternion curves
Ken Shoemake · 1985
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rezero is all you need: Fast convergence at large depth
Thomas Bachlechner, Huanru Henry Majumder, Bodhisattwa Prasad Mao, Garrison W. Cottrell, and Julian McAuley · 2003
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Multi-objective optimization
Kalyanmoy Deb, Karthik Sindhya, and Jussi Hakanen · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Deep hyperspherical learning
Weiyang Liu, Yan-Ming Zhang, Xingguo Li, Zhiding Yu, Bo Dai, Tuo Zhao, and Le Song · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Normface: L2 hypersphere embedding for face verification
Feng Wang, Xiang Xiang, Jian Cheng, and Alan Loddon Yuille · 2017
Earlier work this paper cites.
Fix your classifier: the marginal value of training the last weight layer
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2018
Earlier work this paper cites.
Decoupled networks
Weiyang Liu, Zhen Liu, Zhiding Yu, Bo Dai, Rongmei Lin, Yisen Wang, James M Rehg, and Le Song · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2018
Earlier work this paper cites.
Spherical latent spaces for stable variational autoencoders
Jiacheng Xu and Greg Durrett · 2018
Earlier work this paper cites.
OpenWebText corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Hyperspherical prototype networks
Pascal Mettes, Elise Van der Pol, and Cees Snoek · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Query-key normalization for Transformers
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen · 2020
Cited alongside, same era.
Why do we need weight decay in modern deep learning?
Maksym Andriushchenko, Francesco D’Angelo, Aditya Varre, and Nicolas Flammarion · 2023
Later among the works it cites.
Constrained parameter regularization
Jörg K. H. Franke, Michael Hefenbrock, Gregor Koehler, and Frank Hutter · 2023
Later among the works it cites.
Rotational equilibrium: How weight decay balances learning across neural networks
Atli Kosson, Bettina Messmer, and Martin Jaggi · 2023
Later among the works it cites.
Ilya Loshchilov · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Noam Shazeer · 2020
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola · 2020
Cited alongside, same era.
On layer normalization in the Transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, et al · 2020
Cited alongside, same era.
Adaptive regularization with cubics on manifolds
Naman Agarwal, Nicolas Boumal, Brian Bullins, and Coralia Cartis · 2021
Cited alongside, same era.
Learning by turning: Neural architecture aware optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, and Yisong Yue · 2021
Cited alongside, same era.
Normformer: Improved transformer pretraining with extra normalization
Sam Shleifer, Jason Weston, and Myle Ott · 2021
Cited alongside, same era.
Why can GPT learn in-context? Language models implicitly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei · 2022
Cited alongside, same era.
Later among the works it cites.
Tri Dao and Albert Gu · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, et al · 2024
Closest in time.
Provably optimal memory capacity for modern hopfield models: Tight analysis for transformer-compatible dense associative memories
Jerry Yao-Chieh Hu, Dennis Wu, and Han Liu · 2024
Closest in time.
Analyzing and improving the training dynamics of diffusion models
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine · 2024
Closest in time.
Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, and Johannes von Oswald · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
Scalable optimization in the modular norm
Tim Large, Yang Liu, Jacob Huh, Hyojin Bahng, Phillip Isola, and Jeremy Bernstein · 2025
Closest in time.