Fetching the paper…
Reading the bibliography…
Compressing neural nets is an active research problem, given the large size of state-of-the-art nets for tasks such as object recognition, and the computational limits imposed by mobile devices.
Computers and Intractability: A Guide to the Theory of NP-Completeness
M. R. Garey and D. S. Johnson · 1979
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
S. J. Hanson and L. Y. Pratt · 1989
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations , volume 22 of Applications of Mathematics
A. Benveniste, M. Métivier, and P. Priouret · 1990
Earlier work this paper cites.
Weight discretization paradigm for optical neural networks
E. Fiesler, A. Choudry, and H. J. Caulfield · 1990
Earlier work this paper cites.
Optimal brain damage
Y. LeCun, J. S. Denker, and S. A. Solla · 1990
Earlier work this paper cites.
Generalization by weight-elimination with application to forecasting
A. S. Weigend, D. E. Rumelhart, and B. A. Huberman · 1991
Earlier work this paper cites.
Vector Quantization and Signal Compression
A. Gersho and R. M. Gray · 1992
Earlier work this paper cites.
Simplifying neural networks by soft weight-sharing
S. J. Nowlan and G. E. Hinton · 1992
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
B. Hassibi and D. G. Stork · 1993
Earlier work this paper cites.
Fast neural networks without multipliers
M. Marchesi, G. Orlandi, F. Piazza, and A. Uncini · 1993
Earlier work this paper cites.
Pruning algorithms—a survey
R. Reed · 1993
Earlier work this paper cites.
Multilayer feedforward neural networks with single powers-of-two weights
C. Z. Tang and H. K. Kwan · 1993
Earlier work this paper cites.
Neural Networks for Pattern Recognition
C. M. Bishop · 1995
Earlier work this paper cites.
Sparse approximate solutions to linear systems
B. K. Natarajan · 1995
Earlier work this paper cites.
Optimization of Stochastic Models: The Interface between Simulation and Optimization
G. C. Pflug · 1996
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
D. P. Bertsekas and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Laplacian eigenmaps for dimensionality reduction and data representation
M. Belkin and P. Niyogi · 2003
Earlier work this paper cites.
Stochastic neighbor embedding
G. Hinton and S. T. Roweis · 2003
Cited alongside, same era.
Stochastic Approximation and Recursive Algorithms and Applications
H. J. Kushner and G. G. Yin · 2003
Cited alongside, same era.
Introductory Lectures on Convex Optimization. A Basic Course
Y. Nesterov · 2004
Cited alongside, same era.
Modern Multidimensional Scaling: Theory and Application
I. Borg and P. Groenen · 2005
Cited alongside, same era.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Cited alongside, same era.
Visualizing data using t t -SNE
L. J. P. van der Maaten and G. E. Hinton · 2008
Cited alongside, same era.
The elastic embedding algorithm for dimensionality reduction
Speeding up convolutional neural networks with low rank expansions
M. Jaderberg, A. Vedaldi, and A. Zisserman · 2014
Later among the works it cites.
A fast, universal algorithm to learn parametric nonlinear embeddings
M. Á. Carreira-Perpiñán and M. Vladymyrov · 2015
Later among the works it cites.
Compressing neural networks with the hashing trick
W. Chen, J. Wilson, S. Tyree, K. Weinberger, and Y. Chen · 2015
Later among the works it cites.
BinaryConnect: Training deep neural networks with binary weights during propagations
M. Courbariaux, Y. Bengio, and J.-P. David · 2015
Later among the works it cites.
Compressing deep convolutional networks using vector quantization
Y. Gong, L. Liu, M. Yang, and L. Bourdev · 2015
Later among the works it cites.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Á. Carreira-Perpiñán · 2010
Cited alongside, same era.
Distributed optimization of deeply nested systems
M. Á. Carreira-Perpiñán and W. Wang · 2012
Cited alongside, same era.
Exploiting sparseness in deep neural networks for large vocabulary speech recognition
D. Yu, F. Seide, G. Li, and L. Deng · 2012
Cited alongside, same era.
Ensemble Methods: Foundations and Algorithms
Z.-H. Zhou · 2012
Cited alongside, same era.
Predicting parameters in deep learning
M. Denil, B. Shakibi, L. Dinh, M. Ranzato, and N. de Freitas · 2013
Cited alongside, same era.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arısoy, and B. Ramabhadran · 2013
Cited alongside, same era.
Later among the works it cites.
Learning both weights and connections for efficient neural network
S. Han, J. Pool, J. Tran, and W. Dally · 2015
Later among the works it cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Later among the works it cites.
Tensorizing neural networks
A. Novikov, D. Podoprikhin, A. Osokin, and D. P. Vetrov · 2015
Later among the works it cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding
S. Han, H. Mao, and W. J. Dally · 2016
Later among the works it cites.
Quantized neural networks: Training neural networks with low precision weights and activations
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2016
Later among the works it cites.
F. Li, B. Zhang, and B. Liu · 2016
Later among the works it cites.
XNOR-net: ImageNet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Later among the works it cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Z. Ni, X. Zhou, H. Wen, Y. Wu, and Y. Zou · 2016
Later among the works it cites.
Soft weight-sharing for neural network compression
K. Ullrich, E. Meeds, and M. Welling · 2017
Closest in time.
Trained ternary quantization
C. Zhu, S. Han, H. Mao, and W. J. Dally · 2017
Closest in time.