Fetching the paper…
Reading the bibliography…
How much can pruning algorithms teach us about the fundamentals of learning representations in neural networks? And how much can these fundamentals help while devising new pruning techniques? A lot, it turns out.
What size net gives valid generalization?
Baum, Eric B and Haussler, David · 1989
Earlier work this paper cites.
The cascade-correlation learning architecture
Fahlman, Scott E and Lebiere, Christian · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Yann, Denker, John S, Solla, Sara A, Howard, Richard E, and Jackel, Lawrence D · 1989
Earlier work this paper cites.
Generalization performance of overtrained back-propagation networks
Chauvin, Yves · 1990
Earlier work this paper cites.
Weight quantization in boltzmann machines
Balzer, Wolfgang, Takahashi, Masanobu, Ohta, Jun, and Kyuma, Kazuo · 1991
Earlier work this paper cites.
Fault tolerance of pruned multilayer networks
Segee, Bruce E and Carter, Michael J · 1991
Earlier work this paper cites.
Learning with limited numerical precision using the cascade-correlation algorithm
Hoehfeld, Markus and Fahlman, Scott E · 1992
Cited alongside, same era.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, Babak and Stork, David G · 1993
Cited alongside, same era.
Pruning algorithms-a survey
Reed, Russell · 1993
Cited alongside, same era.
The effects of quantization on multilayer neural networks
Dundar, Gunhan and Rose, Kenneth · 1994
Cited alongside, same era.
MNIST handwritten digit database
LeCun, Yann and Cortes, Corinna · 2010
Cited alongside, same era.
Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua · 2013
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Later among the works it cites.
Neural networks and deep learning
Nielsen, Michael A · 2015
Later among the works it cites.
Reducing communication overhead in distributed learning by an order of magnitude (almost)
Øland, Anders and Raj, Bhiksha · 2015
Later among the works it cites.
Han, Song, Mao, Huizi, and Dally, William J · 2016
Later among the works it cites.
On the compression of recurrent neural networks with an application to lvcsr acoustic modeling for embedded speech recognition
Prabhavalkar, Rohit, Alsharif, Ouais, Bruguier, Antoine, and McGraw, Lan · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Mozer, Michael C and Smolensky, Paul
Cited in the paper.
Using relevance to reduce network size automatically
Mozer, Michael C and Smolensky, Paul
Cited in the paper.