Fetching the paper…
Reading the bibliography…
In an effort to improve generalization in deep learning and automate the process of learning rate scheduling, we propose SALR: a sharpness-aware learning rate update technique designed to recover flat minimizers.
A scale invariant flatness measure for deep network minima
Rangamani, A., Nguyen, N. H., Kumar, A., Phan, D., Chin, S. H., and Tran, T. D. (2019) · 1902
Earlier work this paper cites.
Cyclical stochastic gradient mcmc for bayesian deep learning
Zhang, R., Li, C., Zhang, J., Chen, C., and Wilson, A. G. (2019) · 1902
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V. (2019) · 1905
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P. (1997) · 1997
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Vc dimension of neural networks
Sontag, E. D. (1998) · 1998
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A. (2002) · 2002
Earlier work this paper cites.
On-line learning for very large data sets
Bottou, L. and Le Cun, Y. (2005) · 2005
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O. (2008) · 2008
Earlier work this paper cites.
Rademacher complexity bounds for non-iid processes
Mohri, M. and Rostamizadeh, A. (2009) · 2009
Earlier work this paper cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Zhou, P., Feng, J., Ma, C., Xiong, C., HOI, S., et al. (2020) · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al. (2012) · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R. (2012) · 2012
Earlier work this paper cites.
Densenet: Implementing efficient convnet descriptor pyramids
Iandola, F., Moskewicz, M., Karayev, S., Girshick, R., Darrell, T., and Keutzer, K. (2014) · 2014
Earlier work this paper cites.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
Baldassi, C., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R. (2015) · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Baldassi, C., Borgs, C., Chayes, J. T., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R. (2016) · 2016
Cited alongside, same era.
Distributed deep learning using synchronous stochastic gradient descent
Das, D., Avancha, S., Mudigere, D., Vaidynathan, K., Sridharan, S., Kalamkar, D., Kaul, B., and Dubey, P. (2016) · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2016) · 2016
Cited alongside, same era.
Cyclical learning rates for training neural networks
Smith, L. N. (2017) · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J. (2018) · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
An alternative view: When does sgd escape local minima?
Kleinberg, R., Li, Y., and Yuan, Y. (2018) · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016) · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M. (2017) · 2017
Cited alongside, same era.
Fast rates for empirical risk minimization of strict saddle problems
Gonen, A. and Shalev-Shwartz, S. (2017) · 2017
Cited alongside, same era.
Foundations of machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A. (2018) · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. (2018) · 2018
Later among the works it cites.
Identifying generalization properties in neural networks
Wang, H., Keskar, N. S., Xiong, C., and Socher, R. (2018) · 2018
Later among the works it cites.
Smoothout: Smoothing out sharp minima to improve generalization in deep learning
Wen, W., Wang, Y., Yan, F., Xu, C., Wu, C., Chen, Y., and Li, H. (2018) · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z. (2019) · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R. (2019) · 2019
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2019) · 2019
Later among the works it cites.
A simple baseline for bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G. (2019) · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Shorten, C. and Khoshgoftaar, T. M. (2019) · 2019
Later among the works it cites.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P. (2020) · 2020
Closest in time.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. (2021) · 2021
Closest in time.