Fetching the paper…
Reading the bibliography…
In continual learning, a learner has to keep learning from the data over its whole life time.
Learning feature relevance through step size adaptation in temporal-difference learning
Kearney, A., Veeriah, V., Travnik, J. B., Pilarski, P. M., and Sutton, R. S. (2019) · 1903
Earlier work this paper cites.
Adaptation of learning rate parameters
Sutton, R. S. (1981) · 1981
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, R. S. (1992) · 1992
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M. (1999) · 1999
Earlier work this paper cites.
Online Learning with Adaptive Local Step Sizes
Schraudolph, N. N. (1999) · 1999
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Amari, S., Park, H., and Fukumizu, K. (2000) · 2000
Earlier work this paper cites.
Online Independent Component Analysis With Local Learning Rate Adaptation
Schraudolph, N. N. and Giannakopoulos, X. (2000) · 2000
Earlier work this paper cites.
Markerless tracking of complex human motions from multiple views
Kehl, R. and Van Gool, L. (2006) · 2006
Earlier work this paper cites.
On the role of tracking in stationary environments
Sutton, R. S., Koop, A., and Silver, D. (2007) · 2007
Earlier work this paper cites.
Investigating Experience: Temporal Coherence and Empirical Knowledge Representation
Koop, A. (2008) · 2008
Earlier work this paper cites.
A fast natural newton method
Roux, N. L. and Fitzgibbon, A. W. (2010) · 2010
Cited alongside, same era.
RMSProp: Divide the gradient by a running average of its recent magnitude
Hinton, G., Srivastava, N., and Swersky, K. (2012) · 2012
Cited alongside, same era.
Tuning-free step-size adaptation
Mahmood, A. R., Sutton, R. S., Degris, T., and Pilarski, P. M. (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y. (2013) · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S. (2014) · 2014
Cited alongside, same era.
Metagrad: Multiple learning rates in online learning
van Erven, T. and Koolen, W. M. (2016) · 2016
Later among the works it cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Bernacchia, A., Lengyel, M., and Hennequin, G. (2018) · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D. (2018) · 2018
Later among the works it cites.
Metatrace: Online step-size tuning by meta-gradient descent for reinforcement learning control
Young, K., Wang, B., and Taylor, M. E. (2018) · 2018
Later among the works it cites.
Meta-descent for online, continual prediction
Jacobsen, A., Schlegel, M., Linke, C., Degris, T., White, A., and White, M. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning the learning rate for prediction with expert advice
Koolen, W. M., van Erven, T., and Grünwald, P. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Temporal difference learning methods with automatic step-size adaption for strategic board games: Connect-4 and dots-and-boxes
Thill, M. (2015) · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gómez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N. (2016) · 2016
Cited alongside, same era.
Algorithms for optimization
Kochenderfer, M. J. and Wheeler, T. A. (2019) · 2019
Later among the works it cites.
Wngrad: Learn the learning rate in gradient descent
Wu, X., Ward, R., and Bottou, L. (2020) · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H. P., Hessel, M., Oh, J., Singh, S., and Silver, D. (2020) · 2020
Later among the works it cites.
A history of meta-gradient: Gradient methods for meta-learning
Sutton, R. S. (2022) · 2022
Later among the works it cites.
Natural neural networks
Desjardins, G., Simonyan, K., Pascanu, R., and Kavukcuoglu, K. (2015) · 2079
Closest in time.