Fetching the paper…
Reading the bibliography…
Studying neural network loss landscapes provides insights into the nature of the underlying optimization problems.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Maxout networks
Goodfellow, I., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I., Vinyals, O., and Saxe, A · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
An empirical analysis of deep network loss surfaces
Im, D., M.Tao, and Branson, K · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Snapshot ensembles: Train 1, get M for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J., and Weinberger, K · 2017
Earlier work this paper cites.
On large-batch training for deep learning: generalization gap and sharp minima
Keskar, N., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Cited alongside, same era.
Three factors influencing minima in SGD
Jastrzȩbski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2018
Cited alongside, same era.
Zhang, C., Bengio, S., and Singer, Y · 2019
Later among the works it cites.
The intriguing role of module criticality in the generalization of deep networks
Chatterji, N. S., Neyshabur, B., and Sedghi, H · 2020
Later among the works it cites.
Revisiting “qualitatively characterizing neural network optimization problems”
Frankle, J · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzȩbski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Later among the works it cites.
Bad global minima exist and SGD can reach them
Liu, S., Papailiopoulos, D., and Achlioptas, D · 2020
Later among the works it cites.
On monotonic linear interpolation of neural network parameters
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Cited alongside, same era.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Cited alongside, same era.
Deep ensembles: A loss landscape perspective
Fort, S., Hu, H., and Lakshminarayanan, B · 2019
Cited alongside, same era.
Visualizing and understanding the effectiveness of BERT
Hao, Y., Dong, L., Wei, F., and Xu, K · 2019
Cited alongside, same era.
Partitioned integrators for thermodynamic parameterization of neural networks
Leimkuhler, B., Matthews, C., and Vlaar, T · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J
Cited in the paper.
Lucas, J., Bae, J., Zhang, M., Ba, J., Zemel, R., and Grosse, R · 2020
Later among the works it cites.
Deep learning is singular, and that’s good
Murfet, D., Wei, S., Gong, M., Li, H., Gell-Redman, J., and Quella, T · 2020
Later among the works it cites.
What is being transferred in transfer learning?
Neyshabur, B., Sedghi, H., and Zhang, C · 2020
Later among the works it cites.
Large batch optimization for deep learning: training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Zhang, J., Karimireddy, S., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S · 2020
Later among the works it cites.
Analyzing monotonic linear interpolation in neural network loss landscapes
Lucas, J., Bae, J., Zhang, M., Fort, S., Zemel, R., and Grosse, R · 2021
Closest in time.