Fetching the paper…
Reading the bibliography…
Differential learning rate (DLR), a technique that applies different learning rates to different model parameters, has been widely used in deep learning and achieved empirical success via its various forms.
Minimization of functions having lipschitz continuous first partial derivatives
Larry Armijo · 1966
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Sparse spatial autoregressions
R Kelley Pace and Ronald Barry · 1997
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu, Andrew Y Ng, et al · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Detection of traffic signs in real-world images: The German Traffic Sign Detection Benchmark
Sebastian Houben, Johannes Stallkamp, Jan Salmen, Marc Schlipsing, and Christian Igel · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Layer-specific adaptive learning rates for deep networks
Bharat Singh, Soham De, Yangmuzi Zhang, Thomas Goldstein, and Gavin Taylor · 2015
Earlier work this paper cites.
No more pesky learning rate guessing games
Leslie N Smith · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani · 2019
Cited alongside, same era.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, Yang Zhang, and Jonathan M Cohen · 2019
Cited alongside, same era.
Optimal first-order methods for convex functions with a quadratic upper bound
Baptiste Goujaud, Adrien Taylor, and Aymeric Dieuleveut · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel · 2022
Later among the works it cites.
A dnn optimizer that improves over adabelief by suppression of the adaptive stepsize range
Guoqiang Zhang, Kenta Niwa, and W Bastiaan Kleijn · 2022
Later among the works it cites.
Learning-rate-free learning by d-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
How to fine-tune bert for text classification?
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang · 2019
Cited alongside, same era.
Pytorch image models
Ross Wightman · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
Blockwise adaptivity: Faster training and better generalization in deep learning
Shuai Zheng and James T Kwok · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Efficient first-order methods for convex minimization: a constructive approach
Yoel Drori and Adrien B Taylor · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Adalip: An adaptive learning rate method per layer for stochastic optimization
George Ioannou, Thanos Tagaris, and Andreas Stafylopatis · 2023
Later among the works it cites.
Dog is sgd’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2023
Later among the works it cites.
Prodigy: An expeditiously adaptive parameter-free learner
Konstantin Mishchenko and Aaron Defazio · 2023
Later among the works it cites.
Dept: Decomposed prompt tuning for parameter-efficient fine-tuning
Zhengxiang Shi and Aldo Lipani · 2023
Later among the works it cites.
Sparse neural additive model: Interpretable deep learning with feature selection via group sparsity
Shiyun Xu, Zhiqi Bu, Pratik Chaudhari, and Ian J Barnett · 2023
Later among the works it cites.
Lora-fa: Memory-efficient low-rank adaptation for large language models fine-tuning
Longteng Zhang, Lin Zhang, Shaohuai Shi, Xiaowen Chu, and Bo Li · 2023
Later among the works it cites.
Automatic gradient descent with generalized newton’s method
Zhiqi Bu and Shiyun Xu · 2024
Later among the works it cites.
Parameter-efficient fine-tuning for large models: A comprehensive survey
Zeyu Han, Chao Gao, Jinyang Liu, Sai Qian Zhang, et al · 2024
Later among the works it cites.
Lora+: Efficient low rank adaptation of large models
Soufiane Hayou, Nikhil Ghosh, and Bin Yu · 2024
Later among the works it cites.
Lora-ga: Low-rank adaptation with gradient approximation
Shaowen Wang, Linxi Yu, and Jian Li · 2024
Later among the works it cites.
Lora-pro: Are low-rank adapters properly optimized?
Zhengbo Wang and Jian Liang · 2024
Later among the works it cites.
Why transformers need adam: A hessian perspective
Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo · 2024
Later among the works it cites.