Fetching the paper…
Reading the bibliography…
Deep networks have been revolutionary in improving performance of machine learning and artificial intelligence systems.
Asymptotically optimal classification for multiple tests with empirically observed statistics
Michael Gutman · 1989
Earlier work this paper cites.
Using the functional behavior of neurons for genetic recombination in neural nets training
Nachum Shamir, David Saad, and Emanuel Marom · 1993
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Avrim Blum, Adam Kalai, and John Langford · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
The bellkor solution to the netflix grand prize
Yehuda Koren · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando De Freitas · 2013
Earlier work this paper cites.
Ad click prediction: a view from the trenches
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Earlier work this paper cites.
Reducing overfitting in deep networks by decorrelating representations
Michael Cogswell, Faruk Ahmed, Ross Girshick, Larry Zitnick, and Dhruv Batra · 2015
Cited alongside, same era.
Song Han, Huizi Mao, and William J Dally · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Distillation of deep learning ensembles as a regularisation method
Alan Mosca and George D Magoulas · 2018
Later among the works it cites.
Learning to specialize with knowledge distillation for visual question answering
Jonghwan Mun, Kimin Lee, Jinwoo Shin, and Bohyung Han · 2018
Later among the works it cites.
Deterministic implementations for reproducibility in deep reinforcement learning
Prabhat Nagarajan, Garrett Warnell, and Peter Stone · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E Dahl, and Geoffrey E Hinton · 2018
Cited alongside, same era.
Moonshine: Distilling with cheap convolutions
Elliot J Crowley, Gavin Gray, and Amos J Storkey · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, SeongUk Park, and Nojun Kwak · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
Xu Lan, Xiatian Zhu, and Shaogang Gong · 2018
Cited alongside, same era.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Later among the works it cites.
Zhe Chen, Yuyan Wang, Dong Lin, Derek Cheng, Lichan Hong, Ed Chi, and Claire Cui · 2020
Closest in time.
Analyzing the role of model uncertainty for electronic health records
Michael W Dusenberry, Dustin Tran, Edward Choi, Jonas Kemp, Jeremy Nixon, Ghassen Jerfel, Katherine Heller, and Andrew M Dai · 2020
Closest in time.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen John Maybank, and Dacheng Tao · 2020
Closest in time.
Rafael Muller, Simon Kornblith, and Geoffrey Hinton · 2020
Closest in time.
Understanding and improving knowledge distillation
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H Chi, and Sagar Jain · 2020
Closest in time.