Fetching the paper…
Reading the bibliography…
As deep nets are increasingly used in applications suited for mobile devices, a fundamental dilemma becomes apparent: the trend in deep learning is to grow models to absorb ever-increasing data set sizes; however mobile devices are designed with very little memory and cannot store such large models.
Generalization of back-propagation to recurrent neural networks
Pineda, Fernando J · 1987
Earlier work this paper cites.
Term-weighting approaches in automatic text retrieval
Salton, Gerard and Buckley, Christopher · 1988
Earlier work this paper cites.
Optimal brain damage
LeCun, Yann, Denker, John S, Solla, Sara A, Howard, Richard E, and Jackel, Lawrence D · 1989
Earlier work this paper cites.
Simplifying neural networks by soft weight-sharing
Nowlan, Steven J and Hinton, Geoffrey E · 1992
Earlier work this paper cites.
Neural Networks for Pattern Recognition
Bishop, Christopher M · 1995
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis
Simard, Patrice Y, Steinkraus, Dave, and Platt, John C · 2003
Earlier work this paper cites.
Model compression
Buciluǎ, Cristian, Caruana, Rich, and Niculescu-Mizil, Alexandru · 2006
Earlier work this paper cites.
An empirical evaluation of deep architectures on problems with many factors of variation
Larochelle, Hugo, Erhan, Dumitru, Courville, Aaron C, Bergstra, James, and Bengio, Yoshua · 2007
Earlier work this paper cites.
Sparse feature learning for deep belief networks
Boureau, Y-lan, Cun, Yann L, et al · 2008
Earlier work this paper cites.
Small statistical models by random feature mixing
Ganchev, Kuzman and Dredze, Mark · 2008
Earlier work this paper cites.
Junior: The stanford entry in the urban challenge
Montemerlo, Michael, Becker, Jan, Bhat, Suhrid, Dahlkamp, Hendrik, Dolgov, Dmitri, Ettinger, Scott, Haehnel, Dirk, Hilden, Tim, Hoffmann, Gabe, Huhnke, Burkhard, et al · 2008
Earlier work this paper cites.
Augmented smartphone applications through clone cloud execution
Chun, Byung-Gon and Maniatis, Petros · 2009
Earlier work this paper cites.
Hash kernels for structured data
Shi, Qinfeng, Petterson, James, Dror, Gideon, Langford, John, Smola, Alex, and Vishwanathan, S.V.N · 2009
Earlier work this paper cites.
Feature hashing for large scale multitask learning
Weinberger, Kilian, Dasgupta, Anirban, Langford, John, Smola, Alex, and Attenberg, Josh · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Speech recognition for mobile devices at google
Schuster, Mike · 2010
Earlier work this paper cites.
High-performance neural networks for visual object classification
Cireşan, Dan C, Meier, Ueli, Masci, Jonathan, Gambardella, Luca M, and Schmidhuber, Jürgen · 2011
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, Adam, Ng, Andrew Y, and Lee, Honglak · 2011
Cited alongside, same era.
Torch7: A matlab-like environment for machine learning
Collobert, Ronan, Kavukcuoglu, Koray, and Farabet, Clément · 2011
Cited alongside, same era.
Deep belief networks using discriminative features for phone recognition
Mohamed, Abdel-rahman, Sainath, Tara N, Dahl, George, Ramabhadran, Bhuvana, Hinton, Geoffrey E, and Picheny, Michael A · 2011
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, Geoffrey, Deng, Li, Yu, Dong, Dahl, George E, Mohamed, Abdel-rahman, Jaitly, Navdeep, Senior, Andrew, Vanhoucke, Vincent, Nguyen, Patrick, Sainath, Tara N, et al · 2012
Cited alongside, same era.
Client vs. server architecture: Why google voice search is also much faster than siri @ONLINE, October 2012
Kosner, A.W · 2012
Cited alongside, same era.
Stochastic pooling for regularization of deep convolutional neural networks
Zeiler, Matthew D and Fergus, Rob · 2013
Later among the works it cites.
A reliable effective terascale linear learning system
Agarwal, Alekh, Chapelle, Olivier, Dudík, Miroslav, and Langford, John · 2014
Later among the works it cites.
Do deep nets really need to be deep?
Ba, Jimmy and Caruana, Rich · 2014
Later among the works it cites.
Marginalized denoising auto-encoders for nonlinear representations
Chen, Minmin, Weinberger, Kilian Q., Sha, Fei, and Bengio, Yoshua · 2014
Later among the works it cites.
Low precision storage for deep learning
Courbariaux, M., Bengio, Y., and David, J.-P · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Efficient backprop
LeCun, Yann A, Bottou, Léon, Orr, Genevieve B, and Müller, Klaus-Robert · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P · 2012
Cited alongside, same era.
Deep learning with cots hpc systems
Coates, Adam, Huval, Brody, Wang, Tao, Wu, David, Catanzaro, Bryan, and Andrew, Ng · 2013
Cited alongside, same era.
Predicting parameters in deep learning
Denil, Misha, Shakibi, Babak, Dinh, Laurent, de Freitas, Nando, et al · 2013
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, Jeff, Jia, Yangqing, Vinyals, Oriol, Hoffman, Judy, Zhang, Ning, Tzeng, Eric, and Darrell, Trevor · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Graves, Alex, Mohamed, A-R, and Hinton, Geoffrey · 2013
Cited alongside, same era.
Denton, Emily, Zaremba, Wojciech, Bruna, Joan, LeCun, Yann, and Fergus, Rob · 2014
Later among the works it cites.
Bayesian optimization with inequality constraints
Gardner, Jacob, Kusner, Matt, Weinberger, Kilian, Cunningham, John, et al · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, Ross, Donahue, Jeff, Darrell, Trevor, and Malik, Jitendra · 2014
Later among the works it cites.
Distilling the knowledge in a neural network
Hinton, Geoffrey, Vinyals, Oriol, and Dean, Jeff · 2014
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, Andrej and Fei-Fei, Li · 2014
Later among the works it cites.
Cnn features off-the-shelf: an astounding baseline for recognition
Razavian, Ali Sharif, Azizpour, Hossein, Sullivan, Josephine, and Carlsson, Stefan · 2014
Later among the works it cites.
Learning ordered representations with nested dropout
Rippel, Oren, Gelbart, Michael A, and Adams, Ryan P · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2014
Later among the works it cites.
Visualizing and understanding convolutional networks
Zeiler, Matthew D and Fergus, Rob · 2014
Later among the works it cites.
Deep learning with limited numerical precision
Gupta, Suyog, Agrawal, Ankur, Gopalakrishnan, Kailash, and Narayanan, Pritish · 2015
Closest in time.