Fetching the paper…
Reading the bibliography…
Adam has been shown to outperform gradient descent on large language models by a larger margin than on other tasks, but it is unclear why.
“Dropout: a simple way to prevent neural networks from overfitting”
Nitish Srivastava, Geoffrey. Hinton, Alex Krizhevsky, Ilya Sutskever and Ruslan Salakhutdinov · 1958
Earlier work this paper cites.
“An improved algorithm for neural network classification of imbalanced training sets”
Rangachari Anand, Kishan. Mehrotra, Chilukuri. Mohan and Sanjay Ranka · 1993
Earlier work this paper cites.
“A new algorithm for data compression”
Philip Gage · 1994
Earlier work this paper cites.
“Gradient-Based Learning Applied to Document Recognition”
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick Haffner · 1998
Earlier work this paper cites.
“Building a Large Annotated Corpus of English: The Penn Treebank”
Mitchell. Marcus, Beatrice Santorini and Mary Marcinkiewicz · 2004
Earlier work this paper cites.
“Inequalities on the Lambert function and hyperpower function”
Abdolhossein Hoorfar and Mehdi Hassani · 2008
Earlier work this paper cites.
“ImageNet: A large-scale hierarchical image database”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-Fei · 2009
Earlier work this paper cites.
“Adaptive Subgradient Methods for Online Learning and Stochastic Optimization”
John. Duchi, Elad Hazan and Yoram Singer · 2011
Earlier work this paper cites.
“RMSPROP: Divide the gradient by a running average of its recent magnitude” Lecture notes
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf , 2012
2012
Earlier work this paper cites.
“Zipf’s word frequency law in natural language: A critical review and future directions”
Steven. Piantadosi · 2014
Earlier work this paper cites.
“Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Zipf’s law holds for phrases, not words”
Jake Williams, Paul. Lessard, Suma Desu, Eric. Clark, James. Bagrow, Christopher. Danforth and Peter Sheridan · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey. Hinton · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Neural Machine Translation of Rare Words with Subword Units”
Rico Sennrich, Barry Haddow and Alexandra Birch · 2016
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Cited alongside, same era.
“Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients”
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
“Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates”
Taku Kudo · 2018
Cited alongside, same era.
“Limitations of the empirical Fisher approximation for natural gradient descent”
Frederik Kunstner, Philipp Hennig and Lukas Balles · 2019
Cited alongside, same era.
“Convergence of Gradient Descent on Separable Data”
Mor Nacson, Jason. Lee, Suriya Gunasekar, Pedro Savarese, Nathan Srebro and Daniel Soudry · 2019
Cited alongside, same era.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
“Better plain ViT baselines for ImageNet-1k”
Lucas Beyer, Xiaohua Zhai and Alexander Kolesnikov · 2022
Later among the works it cites.
“Robustness to Unbounded Smoothness of Generalized SignSGD”
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang and Zhenxun Zhuang · 2022
Later among the works it cites.
“The Three Stages of Learning Dynamics in High-dimensional Kernel Methods”
Nikhil Ghosh, Song Mei and Bin Yu · 2022
Later among the works it cites.
“How Does Adaptive Optimization Impact Local Neural Network Geometry?”
Kaiqi Jiang, Dhruv Malik and Yuanzhi Li · 2022
Later among the works it cites.
“Locating and editing factual associations in GPT”
Kevin Meng, David Bau, Alex Andonian and Yonatan Belinkov · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Paszke et al · 2019
Cited alongside, same era.
“Language Models are Unsupervised Multitask Learners” Tech. Report, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Cited alongside, same era.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2020
Cited alongside, same era.
“Does learning require memorization? a short tale about a long tail”
Vitaly Feldman · 2020
Cited alongside, same era.
“Finding the Optimal Vocabulary Size for Neural Machine Translation”
Thamme Gowda and Jonathan May · 2020
Cited alongside, same era.
“Understanding the Difficulty of Training Transformers”
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen and Jiawei Han · 2020
Cited alongside, same era.
“An investigation of why overparameterization exacerbates spurious correlations”
Shiori Sagawa, Aditi Raghunathan, Pang Koh and Percy Liang · 2020
Cited alongside, same era.
“Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank Collapse”
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Singh and Aurélien Lucchi · 2022
Later among the works it cites.
“Vanishing Curvature in Randomly Initialized Deep ReLU Networks”
Antonio Orvieto, Jonas Kohler, Dario Pavllo, Thomas Hofmann and Aurélien Lucchi · 2022
Later among the works it cites.
“How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers”
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit and Lucas Beyer · 2022
Later among the works it cites.
“Linear attention is (maybe) all you need (to understand transformer optimization)”
Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie and Suvrit Sra · 2023
Later among the works it cites.
“Birth of a Transformer: A Memory Viewpoint”
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou and Leon Bottou · 2023
Later among the works it cites.
“A Theoretical Analysis of the Learning Dynamics under Class Imbalance”
Emanuele Francazi, Marco Baity-Jesi and Aurélien Lucchi · 2023
Later among the works it cites.
“Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation”
Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel. Smith and Yee Teh · 2023
Later among the works it cites.
“Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be”
Frederik Kunstner, Jacques Chen, Jonathan Lavington and Mark Schmidt · 2023
Later among the works it cites.
“The quantization model of neural scaling”
Eric. Michaud, Ziming Liu, Uzay Girit and Max Tegmark · 2023
Later among the works it cites.
“Toward Understanding Why Adam Converges Faster Than SGD for Transformers” NeurIPS 2022 Workshop on Optimization for Machine Learning. arXiv/2306.00204, 2023
Yan Pan and Yuanzhi Li · 2023
Later among the works it cites.
“Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization”
Elan Rosenfeld and Andrej Risteski · 2023
Later among the works it cites.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small”
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris and Jacob Steinhardt · 2023
Later among the works it cites.
“Tokenization and the Noiseless Channel”
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan and Ryan Cotterell · 2023
Later among the works it cites.