Fetching the paper…
Reading the bibliography…
Deep neural network models, though very powerful and highly successful, are computationally expensive in terms of space and time.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The concave-convex procedure (CCCP)
A.L. Yuille and A. Rangarajan · 2002
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Training deep and recurrent networks with Hessian-free optimization
J. Martens and I. Sutskever · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude, 2012
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
M.D. Zeiler · 2012
Earlier work this paper cites.
Pylearn2: a machine learning research library
I.J. Goodfellow, D. Warde-Farley, P. Lamblin, V. Dumoulin, M. Mirza, R. Pascanu, J. Bergstra, F. Bastien, and Y. Bengio · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Y. Gong, L. Liu, M. Yang, and L. Bourdev · 2014
Cited alongside, same era.
Proximal Newton-type methods for minimizing composite functions
J.D. Lee, Y. Sun, and M.A. Saunders · 2014
Cited alongside, same era.
Revisiting natural gradient for deep networks
R. Pascanu and Y. Bengio · 2014
Cited alongside, same era.
BinaryConnect: Training deep neural networks with binary weights during propagations
M. Courbariaux, Y. Bengio, and J.P. David · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and F.-F. Li · 2016
Closest in time.
Compression of deep convolutional neural networks for fast and low power mobile applications
Y.-D. Kim, E. Park, S. Yoo, T. Choi, L. Yang, and D. Shin · 2016
Closest in time.
F. Li and B. Liu · 2016
Closest in time.
Neural networks with few multiplications
Z. Lin, M. Courbariaux, R. Memisevic, and Y. Bengio · 2016
Closest in time.
DC proximal Newton for nonconvex optimization problems
A. Rakotomamonjy, R. Flamary, and G. Gasso · 2016
Closest in time.
XNOR-Net: ImageNet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tensorizing neural networks
A. Novikov, D. Podoprikhin, A. Osokin, and D.P. Vetrov · 2015
Cited alongside, same era.
Deep compression: Compressing deep neural network with pruning, trained quantization and Huffman coding
S. Han, H. Mao, and W.J. Dally · 2016
Cited alongside, same era.
Binarized neural networks
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2016
Cited alongside, same era.
Equilibrated adaptive learning rates for non-convex optimization
Y. Dauphin, H. de Vries, and Y. Bengio
Cited in the paper.
RMSprop and equilibrated adaptive learning rates for non-convex optimization
Y. Dauphin, H. de Vries, J. Chung, and Y. Bengio
Cited in the paper.
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Closest in time.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Z. Ni, X. Zhou, H. Wen, Y. Wu, and Y. Zou · 2016
Closest in time.