Fetching the paper…
Reading the bibliography…
The hypothesis that sub-network initializations (lottery) exist within the initializations of over-parameterized networks, which when trained in isolation produce highly generalizable models, has led to crucial insights into network initialization and has enabled efficient inferencing.
Measuring calibration in deep learning
Nixon, J.; Dusenberry, M.; Zhang, L.; Jerfel, G.; and Tran, D. 2019 · 1904
Earlier work this paper cites.
Sparse Transfer Learning via Winning Lottery Tickets
Mehta, R. 2019 · 1905
Earlier work this paper cites.
On Mixup Training: Improved Calibration and Predictive Uncertainty for Deep Neural Networks
Thulasidasan, S.; Chennupati, G.; Bilmes, J.; Bhattacharya, T.; and Michalak, S. 2019 · 1905
Earlier work this paper cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H.; Lan, J.; Liu, R.; and Yosinski, J. 2019 · 1905
Earlier work this paper cites.
Weight Agnostic Neural Networks
Gaier, A.; and Ha, D. 2019 · 1906
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T.; and Zettlemoyer, L. 2019 · 1907
Earlier work this paper cites.
Pseudo-labeling and confirmation bias in deep semi-supervised learning
Arazo, E.; Ortego, D.; Albert, P.; O’Connor, N. E.; and McGuinness, K. 2019 · 1908
Earlier work this paper cites.
Evaluating Lottery Tickets Under Distributional Shifts
Desai, S.; Zhan, H.; and Aly, A. 2019 · 1910
Earlier work this paper cites.
ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
Berthelot, D.; Carlini, N.; Cubuk, E. D.; Kurakin, A.; Sohn, K.; Zhang, H.; and Raffel, C. 2019a · 1911
Earlier work this paper cites.
Rigging the Lottery: Making All Tickets Winners
Evci, U.; Gale, T.; Menick, J.; Castro, P. S.; and Elsen, E. 2019 · 1911
Earlier work this paper cites.
What’s Hidden in a Randomly Weighted Neural Network?
Ramanujan, V.; Wortsman, M.; Kembhavi, A.; Farhadi, A.; and Rastegari, M. 2019 · 1911
Earlier work this paper cites.
The comparison and evaluation of forecasters
DeGroot, M. H.; and Fienberg, S. E. 1983 · 1983
Cited alongside, same era.
Gradient-based learning applied to document recognition
LeCun, Y.; Bottou, L.; Bengio, Y.; and Haffner, P. 1998 · 1998
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M.; and Gupta, S. 2017 · 2002
Cited alongside, same era.
Evaluating predictive uncertainty challenge
Quinonero-Candela, J.; Rasmussen, C. E.; Sinz, F.; Bousquet, O.; and Schölkopf, B. 2005 · 2005
Cited alongside, same era.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T.; and Raftery, A. E. 2007 · 2007
Cited alongside, same era.
Convolutional deep belief networks on cifar-10
Krizhevsky, A.; and Hinton, G. 2010 · 2010
On calibration of modern neural networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Later among the works it cites.
Variational dropout sparsifies deep neural networks
Molchanov, D.; Ashukha, A.; and Vetrov, D. 2017 · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H.; Rasul, K.; and Vollgraf, R. 2017 · 2017
Later among the works it cites.
mixup: Beyond empirical risk minimization
Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017 · 2017
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J.; and Carbin, M. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MNIST handwritten digit database
LeCun, Y.; Cortes, C.; and Burges, C. 2010 · 2010
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Cited alongside, same era.
Uncertainty in deep learning
Gal, Y. 2016 · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P.; Tyree, S.; Karras, T.; Aila, T.; and Kautz, J. 2016 · 2016
Cited alongside, same era.
Mixmatch: A holistic approach to semi-supervised learning
Berthelot, D.; Carlini, N.; Goodfellow, I.; Papernot, N.; Oliver, A.; and Raffel, C. A. 2019b
Cited in the paper.
Adaptive Network Sparsification with Dependent Variational Beta-Bernoulli Dropout
Lee, J.; Kim, S.; Yoon, J.; Lee, H. B.; Yang, E.; and Hwang, S. J. 2018 · 2018
Later among the works it cites.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Gohil, V.; Narayanan, S. D.; and Jain, A. 2019 · 2019
Later among the works it cites.
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Hendrycks, D.; and Dietterich, T. 2019 · 2019
Later among the works it cites.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Morcos, A.; Yu, H.; Paganini, M.; and Tian, Y. 2019 · 2019
Later among the works it cites.
Learning for single-shot confidence calibration in deep neural networks through stochastic inferences
Seo, S.; Seo, P. H.; and Han, B. 2019 · 2019
Later among the works it cites.
Picking Winning Tickets Before Training by Preserving Gradient Flow
Wang, C.; Zhang, G.; and Grosse, R. 2019 · 2019
Later among the works it cites.