Fetching the paper…
Reading the bibliography…
We propose and evaluate new techniques for compressing and speeding up dense matrix multiplications as found in the fully connected and recurrent layers of neural networks for embedded large vocabulary continuous speech recognition (LVCSR).
Summing and nuclear norms in Banach space theory , volume 8
Graham James Oscar Jameson · 1987
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, Sara A Solla, Richard E Howard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Maximum-margin matrix factorization
Nathan Srebro, Jason Rennie, and Tommi S Jaakkola · 2005
Earlier work this paper cites.
Matrix factorization techniques for recommender systems
Yehuda Koren, Robert Bell, and Chris Volinsky · 2009
Earlier work this paper cites.
Improving the speed of neural networks on cpus
Vincent Vanhoucke, Andrew Senior, and Mark Z Mao · 2011
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Nando de Freitas, et al · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Jian Xue, Jinyu Li, and Yifan Gong · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus · 2014
Cited alongside, same era.
Compressing neural networks with the hashing trick
Wenlin Chen, James Wilson, Stephen Tyree, Kilian Weinberger, and Yixin Chen · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and¡ 0.5 MB model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Later among the works it cites.
Learning compact recurrent neural networks
Zhiyun Lu, Vikas Sindhwani, and Tara N Sainath · 2016
Later among the works it cites.
Personalized speech recognition on mobile devices
Ian McGraw, Rohit Prabhavalkar, Raziel Alvarez, Montse Gonzalez Arenas, Kanishka Rao, David Rybach, Ouais Alsharif, Haşim Sak, Alexander Gruenstein, Françoise Beaufays, et al · 2016
Later among the works it cites.
On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition
Rohit Prabhavalkar, Ouais Alsharif, Antoine Bruguier, and Ian McGraw · 2016
Later among the works it cites.
Reexamining low rank matrix factorization for trace norm regularization
Carlo Ciliberto, Dimitris Stamos, and Massimiliano Pontil · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structured transforms for small-footprint deep learning
Vikas Sindhwani, Tara Sainath, and Sanjiv Kumar · 2015
Cited alongside, same era.
On the efficient representation and execution of deep acoustic models
Raziel Alvarez, Rohit Prabhavalkar, and Anton Bakhtin · 2016
Cited alongside, same era.
Deep Speech 2: End-to-end speech recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Cited alongside, same era.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Cited alongside, same era.
Closest in time.
gemmlowp: a small self-contained low-precision GEMM library
Benoit Jacob and Pete Warden · 2017
Closest in time.
Factorization tricks for LSTM networks
Oleksii Kuchaiev and Boris Ginsburg · 2017
Closest in time.
Gram-CTC: Automatic unit selection and target decomposition for sequence labelling
Hairong Liu, Zhenyao Zhu, Xiangang Li, and Sanjeev Satheesh · 2017
Closest in time.
Exploring sparsity in recurrent neural networks
Sharan Narang, Gregory Diamos, Shubho Sengupta, and Erich Elsen · 2017
Closest in time.