On the design of gradient algorithms for digitally implemented adaptive filters
R Gitlin, J Mazo, and M Taylor · 1973
Earlier work this paper cites.
Transient weight misadjustment properties for the finite precision LMS algorithm
S Alexander · 1987
Earlier work this paper cites.
A nonlinear analytical model for the quantized LMS algorithm-the arbitrary step size case
José Carlos M Bermudez and Neil J Bershad · 1996
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Parallel distributed computing using python
Lisandro D Dalcin, Rodrigo R Paz, Pablo A Kler, and Alejandro Cosimo · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
The convex geometry of linear inverse problems
Venkat Chandrasekaran, Benjamin Recht, Pablo A Parrilo, and Alan S Willsky · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Jian Xue, Jinyu Li, and Yifan Gong · 2013
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Original
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Original
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Original
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Mean-normalized stochastic gradient for large-scale deep learning
Simon Wiesler, Alexander Richard, Ralf Schluter, and Hermann Ney · 2014
Earlier work this paper cites.