Fetching the paper…
Reading the bibliography…
We study whether a neural network optimizes to the same, linearly connected minimum under different samples of SGD noise (e.g., random data order and augmentation).
The state of sparsity in deep neural networks, 2019
Gale, T., Elsen, E., and Hooker, S · 1902
Earlier work this paper cites.
Pruning algorithms: A survey
Reed, R · 1993
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training Imagenet in 1 hour, 2017
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Smith, L. N · 2017
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Networks for Imagenet on TPUs, 2018
Google · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Amc: Automl for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S · 2018
Cited alongside, same era.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Mallya, A., Davis, D., and Lazebnik, S · 2018
Cited alongside, same era.
Super-convergence: Very fast training of residual networks using large learning rates
Smith, L. N. and Topin, N · 2018
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J · 2019
Closest in time.
Finding winning tickets with limited (or no) supervision, 2020
Caron, M., Morcos, A., Bojanowski, P., Mairal, J., and Joulin, A · 2020
Closest in time.
Rigging the lottery: Making all tickets winners, 2020
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Closest in time.
The early phase of neural network training
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Closest in time.
An exponential learning rate schedule for deep learning
Li, Z. and Arora, S · 2020
Closest in time.
What’s hidden in a randomly weighted neural network?
Ramanujan, V., Wortsman, M., Kembhavi, A., Farhadi, A., and Rastegari, M · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Smith, S. L. and Le, Q. V · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Cited alongside, same era.
SNIP: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H. S · 2019
Cited alongside, same era.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2019
Cited alongside, same era.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Morcos, A., Yu, H., Paganini, M., and Tian, Y · 2019
Cited alongside, same era.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z · 2019
Cited alongside, same era.
Comparing fine-tuning and rewinding in neural network pruning
Renda, A., Frankle, J., and Carbin, M · 2020
Closest in time.
Winning the lottery with continuous sparsification
Savarese, P., Silva, H., and Maire, M · 2020
Closest in time.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2020
Closest in time.
The sooner the better: Investigating structure of early winning lottery tickets, 2020
Yin, S., Kim, K.-H., Oh, J., Wang, N., Serrano, M., Seo, J.-S., and Choi, J · 2020
Closest in time.
Playing the lottery with rewards and multiple languages: Lottery tickets in rl and nlp
Yu, H., Edunov, S., Tian, Y., and Morcos, A. S · 2020
Closest in time.