Fetching the paper…
Reading the bibliography…
Modern deep neural networks have a large number of parameters, making them very hard to train.
Simulated annealing: theory and applications
Chii-Ruey Hwang · 1988
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi, David G Stork, et al · 1993
Earlier work this paper cites.
A simple weight decay can improve generalization
J Moody, S Hanson, Anders Krogh, and John A Hertz · 1995
Earlier work this paper cites.
Sparsity and incoherence in compressive sampling
Emmanuel Candes and Justin Romberg · 2007
Earlier work this paper cites.
Sparse online learning via truncated gradient
John Langford, Lihong Li, and Tong Zhang · 2009
Earlier work this paper cites.
BVLC caffe model zoo. http://caffe.berkeleyvision.org/model_zoo, 2013
Yangqing Jia · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Truncated power method for sparse eigenvalue problems
Xiao-Tong Yuan and Tong Zhang · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Ng · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Later among the works it cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Later among the works it cites.
Dmytro Mishkin and Jiri Matas · 2015
Later among the works it cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhaoran Wang, Quanquan Gu, Yang Ning, and Han Liu · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al · 2015
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Facebook.ResNet.Torch. https://github.com/facebook/fb.resnet.torch, 2016
Facebook · 2016
Closest in time.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Closest in time.
Training skinny deep neural networks with iterative hard thresholding methods
Xiaojie Jin, Xiaotong Yuan, Jiashi Feng, and Shuicheng Yan · 2016
Closest in time.