Fetching the paper…
Reading the bibliography…
We propose the generalized Newton's method (GeN) -- a Hessian-informed approach that applies to any optimizer such as SGD and Adam, and covers the Newton-Raphson method as a sub-case.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 1. general considerations
Charles George Broyden · 1970
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k** 2)
Yurii Nesterov · 1983
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
Nicol N Schraudolph, Jin Yu, and Simon Günter · 2007
Earlier work this paper cites.
Object detection combining recognition and segmentation
Liming Wang, Jianbo Shi, Gang Song, and I-fan Shen · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu, Andrew Y Ng, et al · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Detection of traffic signs in real-world images: The German Traffic Sign Detection Benchmark
Sebastian Houben, Johannes Stallkamp, Jan Salmen, Marc Schlipsing, and Christian Igel · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
A variance reduced stochastic newton method
Aurelien Lucchi, Brian McWilliams, and Thomas Hofmann · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
No more pesky learning rate guessing games
Leslie N Smith · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
A stochastic quasi-newton method for large-scale optimization
Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Efficient second order online learning by sketching
Haipeng Luo, Alekh Agarwal, Nicolo Cesa-Bianchi, and John Langford · 2016
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
The cost of training nlp models: A concise overview
Or Sharir, Barak Peleg, and Yoav Shoham · 2020
Later among the works it cites.
Fairscale: A general purpose modular pytorch library for high performance and large scale training
FairScale authors · 2021
Later among the works it cites.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
The e2e dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondrej Dusek, and Verena Rieser · 2017
Cited alongside, same era.
Stochastic quasi-newton methods for nonconvex stochastic optimization
Xiao Wang, Shiqian Ma, Donald Goldfarb, and Wei Liu · 2017
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney · 2021
Later among the works it cites.
Automatic, dynamic, and nearly optimal learning rate specification via local quadratic approximation
Yingqiu Zhu, Danyang Huang, Yuan Gao, Rui Wu, Yu Chen, Bo Zhang, and Hansheng Wang · 2021
Later among the works it cites.
Scalable and efficient training of large convolutional neural networks with differential privacy
Zhiqi Bu, Jialin Mao, and Shiyun Xu · 2022
Later among the works it cites.
Measuring the carbon intensity of ai in cloud instances
Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A Smith, Nicole DeCario, and Will Buchanan · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Sketched newton–raphson
Rui Yuan, Alessandro Lazaric, and Robert M Gower · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel · 2022
Later among the works it cites.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al · 2023
Later among the works it cites.
Adler–an efficient hessian-based strategy for adaptive learning rate
Dario Balboni and Davide Bacciu · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Later among the works it cites.
Learning-rate-free learning by d-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Later among the works it cites.
Dog is sgd’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Later among the works it cites.
Dowg unleashed: An efficient universal parameter-free gradient descent method
Ahmed Khaled, Konstantin Mishchenko, and Chi Jin · 2023
Later among the works it cites.
Estimating the carbon footprint of bloom, a 176b parameter language model
Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Qlabgrad: A hyperparameter-free and convergence-guaranteed scheme for deep learning
Minghan Fu and Fang-Xiang Wu · 2024
Closest in time.
Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate
Xiaomeng Jin, Zhiqi Bu, Bhanukiran Vinzamuri, Anil Ramakrishna, Kai-Wei Chang, Volkan Cevher, and Mingyi Hong · 2025
Closest in time.