Fetching the paper…
Reading the bibliography…
The high computational and parameter complexity of neural networks makes their training very slow and difficult to deploy on energy and storage-constrained computing systems.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N. et al. (2014) · 1958
Earlier work this paper cites.
Finite-precision analysis of the pipelined strength-reduced adaptive filter
Goel, M. and Shanbhag, N. (1998) · 1998
Earlier work this paper cites.
VLSI Digital Signal Processing Systems: Design and Implementation
Parhi, K. (2007) · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
Refinement of the upper bounds of the constants in lyapunov’s theorem
Tyurin, I. S. (2010) · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. (2013) · 2013
Earlier work this paper cites.
Maxout networks
Goodfellow, I. J. et al. (2013) · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y. (2013) · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
Deepface: Closing the gap to human-level performance in face verification
Taigman, Y., Yang, M., Ranzato, M., and Wolf, L. (2014) · 2014
Cited alongside, same era.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P. (2015) · 2015
Cited alongside, same era.
The reusable holdout: Preserving validity in adaptive data analysis
Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. (2015) · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P. (2015) · 2015
Cited alongside, same era.
Han, S., Mao, H., and Dally, W. J. (2015) · 2015
Cited alongside, same era.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. (2016) · 2016
Later among the works it cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M. (2017) · 2017
Later among the works it cites.
The high-dimensional geometry of binary neural networks
Anderson, A. G. and Berg, C. P. (2017) · 2017
Later among the works it cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, D., Frean, M., Leary, L., Lewis, J. P., Ma, K. W.-D., and McWilliams, B. (2017) · 2017
Later among the works it cites.
Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Chen, Y.-H., Krishna, T., Emer, J. S., and Sze, V. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y. (2016) · 2016
Cited alongside, same era.
Fixed point quantization of deep convolutional networks
Lin, D., Talathi, S., and Annapureddy, S. (2016) · 2016
Cited alongside, same era.
XNOR-Net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A. (2016) · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Deep k-means: Re-training and parameter sharing with harder cluster assignments for compressing deep convolutions
Wu, J., Wang, Y., Wu, Z., Wang, Z., Veeraraghavan, A., and Lin, Y. (2018a)
Cited in the paper.
Later among the works it cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Later among the works it cites.
Flexpoint: An adaptive numerical format for efficient training of deep neural networks
Köster, U., Webb, T., Wang, X., Nassar, M., Bansal, A. K., Constable, W., Elibol, O., Hall, S., Hornof, L., Khosrowshahi, A., et al. (2017) · 2017
Later among the works it cites.
On the expressive power of deep neural networks
Raghu, M. et al. (2017) · 2017
Later among the works it cites.
Analytical guarantees on numerical precision of deep neural networks
Sakr, C., Kim, Y., and Shanbhag, N. (2017) · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H. (2017) · 2017
Later among the works it cites.