Fetching the paper…
Reading the bibliography…
Ensembles of machine learning models yield improved system performance as well as robust and interpretable uncertainty estimates; however, their inference costs may often be prohibitively high.
“Ensemble methods in machine learning,”
Thomas G. Dietterich, · 2000
Earlier work this paper cites.
“Estimating a dirichlet distribution,” 2000
Thomas Minka, · 2000
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“ImageNet: A Large-Scale Hierarchical Image Database,”
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, · 2009
Earlier work this paper cites.
“Distilling the knowledge in a neural network,” 2015,
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,”
B. Lakshminarayanan, A. Pritzel, and C. Blundell, · 2017
Earlier work this paper cites.
“Attention is all you need,” 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Understanding Measures of Uncertainty for Adversarial Example Detection,”
L. Smith and Y. Gal, · 2018
Cited alongside, same era.
“Accurate, large minibatch sgd: Training imagenet in 1 hour,” 2018
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2018
Cited alongside, same era.
“Albumentations: fast and flexible image augmentations,”
A. Buslaev, A. Parinov, E. Khvedchenya, V. I. Iglovikov, and A. A. Kalinin, · 2018
Cited alongside, same era.
“Scaling neural machine translation,”
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli, · 2018
Cited alongside, same era.
“A call for clarity in reporting BLEU scores,”
Matt Post, · 2018
Cited alongside, same era.
“Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift,”
“Benchmarking neural network robustness to common corruptions and perturbations,”
Dan Hendrycks and Thomas Dietterich, · 2019
Later among the works it cites.
“Pitfalls of in-domain uncertainty estimation and ensembling in deep learning,”
Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov, · 2020
Later among the works it cites.
“Ensemble distribution distillation,”
Andrey Malinin, Bruno Mlodozeniec, and Mark JF Gales, · 2020
Later among the works it cites.
“Hydra: Preserving ensemble diversity for model distillation,” 2020
Linh Tran, Bastiaan S. Veeling, Kevin Roth, Jakub Świątkowski, Joshua V. Dillon, Jasper Snoek, Stephan Mandt, Tim Salimans, Sebastian Nowozin, and Rodolphe Jenatton, · 2020
Later among the works it cites.
“Ensemble approaches for uncertainty in spoken language assessment,”
Xixin Wu, Kate M Knill, Mark JF Gales, and Andrey Malinin, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V Dillon, Balaji Lakshminarayanan, and Jasper Snoek, · 2019
Cited alongside, same era.
Uncertainty Estimation in Deep Learning with application to Spoken Language Assessment
Andrey Malinin, · 2019
Cited alongside, same era.
“Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning,” 2019
Andreas Kirsch, Joost van Amersfoort, and Yarin Gal, · 2019
Cited alongside, same era.
“Reverse kl-divergence training of prior networks: Improved uncertainty and adversarial robustness,”
Andrey Malinin and Mark JF Gales, · 2019
Cited alongside, same era.
“Fixing the train-test resolution discrepancy,”
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Hervé Jégou, · 2019
Cited alongside, same era.
“Residual learning without normalization via better initialization,”
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma, · 2019
Cited alongside, same era.
Andrey Malinin, Sergey Chervontsev, Ivan Provilkov, and Mark Gales, · 2020
Later among the works it cites.
“Uncertainty in structured prediction,”
Andrey Malinin and Mark Gales, · 2020
Later among the works it cites.
“Rezero is all you need: Fast convergence at large depth,”
Thomas Bachlechner, Huanru Henry Majumder, Bodhisattwa Prasad Mao, Garrison W. Cottrell, and Julian McAuley, · 2020
Later among the works it cites.
“The many faces of robustness: A critical analysis of out-of-distribution generalization,”
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer, · 2020
Later among the works it cites.
“Natural adversarial examples,”
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song, · 2021
Closest in time.