Fetching the paper…
Reading the bibliography…
The Transformer architecture has inarguably revolutionized deep learning, overtaking classical architectures like multi-layer perceptrons (MLPs) and convolutional neural networks (CNNs).
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Solutions to some functional equations and their applications to characterization of probability distributions
CG Khatri and C Radhakrishna Rao · 1968
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Matrix results on the khatri-rao and tracy-singh products
Shuangzhe Liu · 1999
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Hand-waving and interpretive dance: an introductory course on tensor networks
Jacob C Bridgeman and Christopher T Chubb · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Training tips for the transformer model
Martin Popel and Ondřej Bojar · 2018
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2020
Later among the works it cites.
Traces of class/cross-class structure pervade deep learning spectra
Vardan Papyan · 2020
Later among the works it cites.
Woodfisher: Efficient second-order approximation for neural network compression
Sidak Pal Singh and Dan Alistarh · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Cited alongside, same era.
Limitations of the empirical fisher approximation for natural gradient descent
Frederik Kunstner, Philipp Hennig, and Lukas Balles · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Matrix differential calculus with applications in statistics and econometrics
Jan R Magnus and Heinz Neudecker · 2019
Cited alongside, same era.
Later among the works it cites.
Analytic insights into structure and rank of neural network hessian maps
Sidak Pal Singh, Gregor Bachmann, and Thomas Hofmann · 2021
Later among the works it cites.
Escaping the gradient vanishing: Periodic alternatives of softmax in attention mechanism
Shulun Wang, Feng Liu, and Bin Liu · 2021
Later among the works it cites.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Later among the works it cites.
Toward understanding why adam converges faster than sgd for transformers
Yan Pan and Yuanzhi Li · 2022
Later among the works it cites.
How do vision transformers work?
Namuk Park and Songkuk Kim · 2022
Later among the works it cites.
The hessian perspective into the nature of convolutional neural networks
Sidak Pal Singh, Thomas Hofmann, and Bernhard Schölkopf · 2023
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, and Suvrit Sra · 2024
Closest in time.
Understanding addition in transformers
Philip Quirke and Fazl Barez · 2024
Closest in time.
Why transformers need adam: A hessian perspective
Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo · 2024
Closest in time.
Position: Curvature matrices should be democratized via linear operators
Felix Dangel, Runa Eschenhagen, Weronika Ormaniec, Andres Fernandez, Lukas Tatzel, and Agustinus Kristiadi · 2025
Closest in time.
Adam-mini: Use fewer learning rates to gain more
Yushun Zhang, Congliang Chen, Ziniu Li, Tian Ding, Chenwei Wu, Diederik P Kingma, Yinyu Ye, Zhi-Quan Luo, and Ruoyu Sun · 2025
Closest in time.