Fetching the paper…
Reading the bibliography…
Yes, they do.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Antonio Torralba, Robert Fergus, and William T. Freeman · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2010
Earlier work this paper cites.
Conversational speech transcription using context-dependent deep neural networks
Frank Seide, Gang Li, and Dong Yu · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian J. Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Big neural networks waste capacity
Yann N. Dauphin and Yoshua Bengio · 2013
Earlier work this paper cites.
Fastfood-computing hilbert space expansions in loglinear time
Quoc Le, Tamás Sarlós, and Alexander Smola · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Cited alongside, same era.
Understanding deep architectures using a recursive convolutional network
David Eigen, Jason Rolfe, Rob Fergus, and Yann LeCun · 2014
Cited alongside, same era.
Learning small-size dnn with output-distribution-based criteria
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Transferring knowledge from a RNN to a DNN
William Chan, Nan Rosemary Ke, and Ian Laner · 2015
Zero-bias autoencoders and the benefits of co-adapting features
Roland Memisevic, Kishore Konda, and David Krueger · 2015
Later among the works it cites.
FitNets: Hints for thin deep nets
Adriana Romero, Ballas Nicolas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Later among the works it cites.
Scalable bayesian optimization using deep neural networks
Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Md Patwary, Mostofa Ali, Ryan P Adams, et al · 2015
Later among the works it cites.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Juergen Schmidhuber · 2015
Later among the works it cites.
Convolutional rectifier networks as generalized tensor decompositions
Nadav Cohen and Amnon Shashua · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scheduled denoising autoencoders
Krzysztof J. Geras and Charles Sutton · 2015
Cited alongside, same era.
Krzysztof J. Geras, Abdel-rahman Mohamed, Rich Caruana, Gregor Urban, Shengjie Wang, Ozlem Aslan, Matthai Philipose, Matthew Richardson, and Charles Sutton · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Shiyu Liang and R Srikant · 2016
Closest in time.
How far can we go without convolution: Improving fully-connected networks
Zhouhan Lin, Roland Memisevic, Shaoqing Ren, and Kishore Konda · 2016
Closest in time.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2016
Closest in time.
Policy distillation
Andrei A. Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Closest in time.
Regularizing neural networks by penalizing output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey Hinton · 2017
Closest in time.