Fetching the paper…
Reading the bibliography…
The delta-bar-delta algorithm is recognized as a learning rate adaptation technique that enhances the convergence speed of the training process in optimization by dynamically scheduling the learning rate based on the difference between the current and previous weight updates.
Adaptive switching circuits
Bernard Widrow, Marcian E Hoff, et al · 1960
Earlier work this paper cites.
Goal seeking components for adaptive intelligence: An initial assessment
Andrew G Barto and Richard S Sutton · 1981
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker, Yann Le Cun, et al · 1988
Earlier work this paper cites.
An empirical study of learning speed in back-propagation networks
Scott E Fahlman et al · 1988
Earlier work this paper cites.
Increased rates of convergence through learning rate adaptation
Robert A Jacobs · 1988
Earlier work this paper cites.
Back-propagation heuristics: a study of the extended delta-bar-delta algorithm
Ali A Minai and Ronald D Williams · 1990
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard S Sutton · 1992
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
A perspective view and survey of meta-learning
Ricardo Vilalta and Youssef Drissi · 2002
Earlier work this paper cites.
Meta-learning in reinforcement learning
Nicolas Schweighofer and Kenji Doya · 2003
Earlier work this paper cites.
An enhanced version of delta-bar-delta algorithm
Mohammed A Otair, Jordan Irbed, and Walid A Salameh · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun, Corinna Cortes, and Christopher J.C. Burges · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization
Sashank J Reddi, Suvrit Sra, Barnabas Poczos, and Alexander J Smola · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Earlier work this paper cites.
Exponential decay sine wave learning rate for fast deep neural network training
Wangpeng An, Haoqian Wang, Yulun Zhang, and Qionghai Dai · 2017
Earlier work this paper cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Earlier work this paper cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Earlier work this paper cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Earlier work this paper cites.
Learning non-overlapping convolutional neural networks with multiple kernels
Kai Zhong, Zhao Song, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Joaquin Vanschoren · 2018
Earlier work this paper cites.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Zhiming Zhou, Qingru Zhang, Guansong Lu, Hongwei Wang, Weinan Zhang, and Yong Yu · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Memory-efficient adaptive optimization for large-scale learning
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2021
Later among the works it cites.
Fl-ntk: A neural tangent kernel-based framework for federated learning analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Later among the works it cites.
Autolrs: Automatic learning-rate schedule by bayesian optimization on the fly
Yuchen Jin, Tianyi Zhou, Liangyu Zhao, Yibo Zhu, Chuanxiong Guo, Marco Canini, and Arvind Krishnamurthy · 2021
Later among the works it cites.
How to decay your learning rate
Aitor Lewkowycz · 2021
Later among the works it cites.
Autolr: Layer-wise pruning and auto-tuning of learning rates in fine-tuning of deep networks
Youngmin Ro and Jin Young Choi · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Cited alongside, same era.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Aditya Sharad Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Cited alongside, same era.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, Yang Zhang, and Jonathan M Cohen · 2019
Cited alongside, same era.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Rong Ge, Sham M Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Cited alongside, same era.
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Later among the works it cites.
Strength of minibatch noise in sgd
Liu Ziyin, Kangqiao Liu, Takashi Mori, and Masahito Ueda · 2021
Later among the works it cites.
Sgd can converge to local maxima
Liu Ziyin, Botao Li, James B Simon, and Masahito Ueda · 2021
Later among the works it cites.
Step-size adaptation using exponentiated gradient updates
Ehsan Amid, Rohan Anil, Christopher Fifty, and Manfred K Warmuth · 2022
Later among the works it cites.
Bounding the width of neural networks via coupled initialization a worst case analysis
Alexander Munteanu, Simon Omlor, Zhao Song, and David Woodruff · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2023
Closest in time.
A survey of meta-reinforcement learning
Jacob Beck, Risto Vuorio, Evan Zheran Liu, Zheng Xiong, Luisa Zintgraf, Chelsea Finn, and Shimon Whiteson · 2023
Closest in time.
Learning with limited samples: Meta-learning and applications to communication systems
Lisha Chen, Sharu Theresa Jose, Ivana Nikoloska, Sangwoo Park, Tianyi Chen, Osvaldo Simeone, et al · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
Fine-tune language models to approximate unbiased in-context learning
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Closest in time.
How to protect copyright data in optimization of large language models?
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Learning-rate-free learning by d-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Closest in time.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Jinbo Hu, Yin Cao, Ming Wu, Feiran Yang, Ziying Yu, Wenwu Wang, Mark D Plumbley, and Jun Yang · 2023
Closest in time.
Improving adaptation/learning transients using a dynamic adaptation gain/learning rate-theoretical and experimental results
Ioan Doré Landau, Tudor-Bogdan Airimitoaie, Bernard Vau, and Gabriel Buche · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Shuai Li, Zhao Song, Yu Xia, Tong Yu, and Tianyi Zhou · 2023
Closest in time.
Federated adversarial learning: A framework with convergence analysis
Xiaoxiao Li, Zhao Song, and Jiaming Yang · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Closest in time.
Meta-learning a cross-lingual manifold for semantic parsing
Tom Sherborne and Mirella Lapata · 2023
Closest in time.
A unified scheme of resnet and softmax
Zhao Song, Weixin Wang, and Junze Yin · 2023
Closest in time.
Efficient asynchronize stochastic gradient algorithm with structured data
Zhao Song and Mingquan Ye · 2023
Closest in time.
Adaptive compositional continual meta-learning
Bin Wu, Jinyuan Fang, Xiangxiang Zeng, Shangsong Liang, and Qiang Zhang · 2023
Closest in time.
Infoprompt: Information-theoretic soft prompt tuning for natural language understanding
Junda Wu, Tong Yu, Rui Wang, Zhao Song, Ruiyi Zhang, Handong Zhao, Chaochao Lu, Shuai Li, and Ricardo Henao · 2023
Closest in time.
Xin Yuan, Pedro Savarese, and Michael Maire · 2023
Closest in time.