Fetching the paper…
Reading the bibliography…
Ensembles of neural networks are known to be much more robust and accurate than individual networks.
Neural network ensembles
Lars Kai Hansen and Peter Salamon · 1990
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
Anders Krogh, Jesper Vedelsby, et al · 1995
Earlier work this paper cites.
Fast committee learning: Preliminary results
A Swann and N Allinson · 1998
Earlier work this paper cites.
Ensemble selection from libraries of models
Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew, and Alex Ksikes · 2004
Earlier work this paper cites.
Model compression
Cristian Bucilu¨£, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
Ronan Collobert, Koray Kavukcuoglu, and Clément Farabet · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning, 2011
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Convolutional neural networks applied to house numbers digit classification
Pierre Sermanet, Soumith Chintala, and Yann LeCun · 2012
Earlier work this paper cites.
Maxout networks
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus · 2013
Cited alongside, same era.
Horizontal and vertical ensemble with deep representation for classification
Jingjing Xie, Bing Xu, and Zhang Chuang · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Deeply-supervised nets
Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu · 2015
Later among the works it cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Later among the works it cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Later among the works it cites.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron Courville, et al · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Later among the works it cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2016
Later among the works it cites.
Fractalnet: Ultra-deep neural networks without residuals
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2016
Later among the works it cites.
Sgdr: Stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Later among the works it cites.
Boosted convolutional neural networks
Mohammad Moghimi, Mohammad Saberian, Jian Yang, Li-Jia Li, Nuno Vasconcelos, and Serge Belongie · 2016
Later among the works it cites.
Edinburgh neural machine translation systems for wmt 16
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Later among the works it cites.
Swapout: Learning an ensemble of deep architectures
Saurabh Singh, Derek Hoiem, and David Forsyth · 2016
Later among the works it cites.
No more pesky learning rate guessing games
Leslie N. Smith · 2016
Later among the works it cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.