Fetching the paper…
Reading the bibliography…
One aim shared by multiple settings, such as continual learning or transfer learning, is to leverage previously acquired knowledge to converge faster on the current task.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J. N., Pascanu, R., Rabinowitz, N. C., Veness, J., and et. al · 2017
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Mallya, A. and Lazebnik, S · 2017
Earlier work this paper cites.
Riemannian walk for incremental learning: Understanding forgetting and intransigence
Chaudhry, A., Dokania, P. K., Ajanthan, T., and Torr, P. H. S · 2018
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2019
Cited alongside, same era.
On the difficulty of warm-starting neural network training
Ash, J. T. and Adams, R. P · 2019
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Cited alongside, same era.
Rethinking imagenet pre-training
He, K., Girshick, R., and Dollár, P · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M · 2019
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Later among the works it cites.
Continual backprop: Stochastic gradient descent with persistent randomness
Dohare, S., Sutton, R., and Mahmood, A. R · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Golatkar, A. S., Achille, A., and Soatto, S · 2019
Cited alongside, same era.
Transient non- stationarity and generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Böhmer, W., and Whiteson, S · 2021
Closest in time.