Fetching the paper…
Reading the bibliography…
Ensembles over neural network weights trained from different random initialization, known as deep ensembles, achieve state-of-the-art accuracy and calibration.
Verification of forecasts expressed in terms of probability
G. W. Brier · 1950
Earlier work this paper cites.
Neural network ensembles
L. K. Hansen and P. Salamon · 1990
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel · 1990
Earlier work this paper cites.
A statistical approach to learning and generalization in layered neural networks
E. Levin, N. Tishby, and S. A. Solla · 1990
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
S. Geman, E. Bienenstock, and R. Doursat · 1992
Earlier work this paper cites.
A primer in game theory
R. Gibbons et al · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
J. Schmidhuber · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
G. E. Hinton and D. Van Camp · 1993
Earlier work this paper cites.
A ‘self-referential’weight matrix
J. Schmidhuber · 1993
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
A. Krogh and J. Vedelsby · 1995
Earlier work this paper cites.
Ensemble learning and evidence maximization
D. J. MacKay et al · 1995
Earlier work this paper cites.
Bayesian learning for neural networks
R. M. Neal · 1995
Earlier work this paper cites.
Online algorithms and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Popular ensemble methods: An empirical study
D. Opitz and R. Maclin · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
Ensemble selection from libraries of models
R. Caruana, A. Niculescu-Mizil, G. Crew, and A. Ksikes · 2004
Earlier work this paper cites.
Getting the most out of ensemble selection
R. Caruana, A. Munson, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
An overview of bilevel optimization
B. Colson, P. Marcotte, and G. Savard · 2007
Earlier work this paper cites.
The discovery of structural form
C. Kemp and J. B. Tenenbaum · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Learning the structure of deep sparse graphical models
R. Adams, H. Wallach, and Z. Ghahramani · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
J. Bergstra, R. Bardenet, Y. Bengio, B. Kégl, et al · 2011
Earlier work this paper cites.
Towards fully autonomous driving: Systems and algorithms
J. Levinson, J. Askeland, J. Becker, J. Dolson, D. Held, S. Kammel, J. Z. Kolter, D. Langer, O. Pink, V. Pratt, et al · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
Structure discovery in nonparametric regression through compositional kernel search
D. Duvenaud, J. Lloyd, R. Grosse, J. Tenenbaum, and G. Zoubin · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Weight uncertainty in neural network
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Efficient and robust automated machine learning
M. Feurer, A. Klein, K. Eggensperger, J. Springenberg, M. Blum, and F. Hutter · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Practical automated machine learning for the automl challenge 2018
M. Feurer, K. Eggensperger, S. Falkner, M. Lindauer, and F. Hutter · 2018
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
D. Hendrycks and T. Dietterich · 2018
Later among the works it cites.
Stochastic hyperparameter optimization through hypernetworks
J. Lorraine and D. Duvenaud · 2018
Later among the works it cites.
Self-tuning networks: Bilevel optimization of hyperparameters using structured best-response functions
M. Mackay, P. Vicol, J. Lorraine, D. Duvenaud, and R. Grosse · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum · 2015
Cited alongside, same era.
Why m heads are better than one: Training a diverse ensemble of deep networks
S. Lee, S. Purushwalkam, M. Cogswell, D. Crandall, and D. Batra · 2015
Cited alongside, same era.
Obtaining well calibrated probabilities using bayesian binning
M. P. Naeini, G. Cooper, and M. Hauskrecht · 2015
Cited alongside, same era.
Scalable Bayesian optimization using deep neural networks
J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. Patwary, M. Prabhat, and R. Adams · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
D. Ha, A. Dai, and Q. V. Le · 2016
Cited alongside, same era.
Flipout: Efficient pseudo-independent weight perturbations on mini-batches
Y. Wen, P. Vicol, J. Ba, D. Tran, and R. Grosse · 2018
Later among the works it cites.
Adjustable real-time style transfer
M. Babaeizadeh and G. Ghiasi · 2019
Later among the works it cites.
Hyperparameter optimization
M. Feurer and F. Hutter · 2019
Later among the works it cites.
Deep ensembles: A loss landscape perspective
S. Fort, H. Hu, and B. Lakshminarayanan · 2019
Later among the works it cites.
Why relu networks yield high-confidence predictions far away from the training data and how to mitigate the problem
M. Hein, M. Andriushchenko, and J. Bitterwolf · 2019
Later among the works it cites.
Measuring calibration in deep learning
J. Nixon, M. W. Dusenberry, L. Zhang, G. Jerfel, and D. Tran · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
J. Snoek, Y. Ovadia, E. Fertig, B. Lakshminarayanan, S. Nowozin, D. Sculley, J. Dillon, J. Ren, and Z. Nado · 2019
Later among the works it cites.
Bayesian Layers: A module for neural network uncertainty
D. Tran, M. W. Dusenberry, D. Hafner, and M. van der Wilk · 2019
Later among the works it cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Y. Wen, D. Tran, and J. Ba · 2019
Later among the works it cites.
Cyclical stochastic gradient mcmc for bayesian deep learning
R. Zhang, C. Li, J. Zhang, C. Chen, and A. G. Wilson · 2019
Later among the works it cites.
Depth uncertainty in neural networks
J. Antorán, J. U. Allingham, and J. M. Hernández-Lobato · 2020
Closest in time.
You only train once: Loss-conditional training of deep networks
A. Dosovitskiy and J. Djolonga · 2020
Closest in time.
Efficient and scalable bayesian neural nets with rank-1 factors
M. W. Dusenberry, G. Jerfel, Y. Wen, Y.-a. Ma, J. Snoek, K. Heller, B. Lakshminarayanan, and D. Tran · 2020
Closest in time.
Evaluating scalable bayesian deep learning methods for robust computer vision
F. K. Gustafsson, M. Danelljan, and T. B. Schon · 2020
Closest in time.
Large-scale dna-based phenotypic recording and deep learning enable highly accurate sequence-function mapping
S. Höllerer, L. Papaxanthos, A. C. Gumpinger, K. Fischer, C. Beisel, K. Borgwardt, Y. Benenson, and M. Jeschek · 2020
Closest in time.
A deep learning system for differential diagnosis of skin diseases
Y. Liu, A. Jain, C. Eng, D. H. Way, K. Lee, P. Bui, K. Kanada, G. de Oliveira Marinho, J. Gallegos, S. Gabriele, et al · 2020
Closest in time.
Optimized generic feature learning for few-shot classification across domains
T. Saikia, T. Brox, and C. Schmid · 2020
Closest in time.
How good is the bayes posterior in deep neural networks really?
F. Wenzel, K. Roth, B. S. Veeling, J. Świątkowski, L. Tran, S. Mandt, J. Snoek, T. Salimans, R. Jenatton, and S. Nowozin · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
A. G. Wilson and P. Izmailov · 2020
Closest in time.
Neural ensemble search for performant and calibrated predictions
S. Zaidi, A. Zela, T. Elsken, C. Holmes, F. Hutter, and Y. W. Teh · 2020
Closest in time.