Fetching the paper…
Reading the bibliography…
In automatic speech recognition (ASR), model pruning is a widely adopted technique that reduces model size and latency to deploy neural network models on edge devices with resource constraints.
“Optimal brain damage,”
Y. LeCun, J. S. Denker, and S. A. Solla, · 1990
Earlier work this paper cites.
“Second order derivatives for network pruning: Optimal brain surgeon,”
B. Hassibi and D. G. Stork, · 1993
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li, · 2009
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Regularization of neural networks using dropconnect,”
L. Wan, M. Zeiler, S. Zhang, Y. LeCun, and R. Fergus, · 2013
Earlier work this paper cites.
“Learning both weights and connections for efficient neural network,”
S. Han, J. Pool, J. Tran, and W. Dally, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,” 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al., · 2015
Earlier work this paper cites.
“Dynamic network surgery for efficient DNNs,”
Y. Guo, A. Yao, and Y. Chen, · 2016
Earlier work this paper cites.
“Compression of neural machine translation models via pruning,”
A. See, M. Luong, and C. D. Manning, · 2016
Earlier work this paper cites.
“ThiNet: A filter level pruning method for deep neural network compression,”
J. Luo, J. Wu, and W. Lin, · 2017
Earlier work this paper cites.
“Shallow-fusion end-to-end contextual biasing,”
D. Zhao, T. N. Sainath, D. Rybach, P. Rondon, D. Bhatia, B. Li, and R. Pang, · 2017
Earlier work this paper cites.
“Generated of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Cited alongside, same era.
“Reducing the computational complexity for whole word models,”
H. Soltau, H. Liao, and H. Sak, · 2017
Cited alongside, same era.
“Block-sparse recurrent neural networks,”
S. Narang, E. Undersander, and G. Diamos, · 2017
Cited alongside, same era.
“Understanding deep learning requires rethinking generalization,” 2017
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, · 2017
Cited alongside, same era.
“A closer look at memorization in deep networks,”
D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien, · 2017
Cited alongside, same era.
“Deconstructing lottery tickets: Zeros, signs, and the supermask,”
H. Zhou, J. Lan, R. Liu, and J. Yosinski, · 2019
Later among the works it cites.
“Slimmable neural networks,”
J. Yu, L. Yang, N. Xu, J. Yang, and T. Huang, · 2019
Later among the works it cites.
“Universally slimmable networks and improved training techniques,”
J. Yu and T. Huang, · 2019
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, et al., · 2019
Later among the works it cites.
“Rethinking the value of network pruning,”
Z. Liu, M. Sun, T. Zhou, G. Huang, and T. Darrell, · 2019
Later among the works it cites.
“SNIP: Single-shot network pruning based on connection sensitivity,”
N. Lee, T. Ajanthan, and P. Torr, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“To prune, or not to prune: Exploring the efficacy of pruning for model compression,” 2018
M. H. Zhu and S. Gupta, · 2018
Cited alongside, same era.
“AI benchmark: Running deep neural networks on android smartphones,”
A. Ignatov, R. Timofte, W. Chou, K. Wang, M. Wu, T. Hartley, and L. Van Gool, · 2018
Cited alongside, same era.
“Soft filter pruning for accelerating deep convolutional neural networks,”
Y. He, G. Kang, X. Dong, Y. Fu, and Y. Yang, · 2018
Cited alongside, same era.
““Learning-compression” algorithms for neural net pruning,”
M. A. Carreira-Perpinán and Y. Idelbayev, · 2018
Cited alongside, same era.
“Dynamic deep neural networks: Optimizing accuracy-efficiency trade-offs by selective execution,”
L. Liu and J. Deng, · 2018
Cited alongside, same era.
“Multi-scale dense networks for resource efficient image classification,”
G. Huang, D. Chen, T. Li, F. Wu, L. V. Maaten, and K. Weinberger, · 2018
Cited alongside, same era.
“Optimizing speech recognition for the edge,”
Y. Shangguan, J. Li, L. Qiao, R. Alvarez, and I. McGraw, · 2019
Cited alongside, same era.
“The lottery ticket hypothesis: Finding sparse, trainable neural networks,”
J. Frankle and M. Carbin, · 2019
Later among the works it cites.
“Stabilizing the lottery ticket hypothesis,” 2019
J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin, · 2019
Later among the works it cites.
“Network slimming by slimmable networks: Towards one-shot architecture search for channel numbers,”
J. Yu and T. Huang, · 2019
Later among the works it cites.
“What’s hidden in a randomly weighted neural network?,”
V. Ramanujan, M. Wortsman, A. Kembhavi, A. Farhadi, and M. Rastegari, · 2020
Closest in time.
“BigNAS: Scaling up neural architecture search with big single-stage models,”
J. Yu, P. Jin, H. Liu, G. Bender, P. Kindermans, M. Tan, T. Huang, X. Song, R. Pang, and Q. Le, · 2020
Closest in time.
“Efficient knowledge distillation for rnn-transducer models,” 2020
S. Panchapagesan, D. S. Park, C. Chiu, Y. Shangguan, Q. Liang, and A. Gruenstein, · 2020
Closest in time.