Fetching the paper…
Reading the bibliography…
Structured matrices, such as those derived from Kronecker products (KP), are effective at compressing neural networks, but can lead to unacceptable accuracy loss when applied to large models.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 1902
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 1902
Earlier work this paper cites.
Measuring scheduling efficiency of rnns for NLP applications
Thakker, U., Dasika, G., Beu, J. G., and Mattina, M · 1904
Earlier work this paper cites.
Compressing rnns for iot devices by 15-38x using kronecker products
Thakker, U., Beu, J. G., Gope, D., Zhou, C., Fedorov, I., Dasika, G., and Mattina, M · 1906
Earlier work this paper cites.
Pushing the limits of rnn compression
Thakker, U., Fedorov, I., Beu, J. G., Gope, D., Zhou, C., Dasika, G., and Mattina, M · 1910
Earlier work this paper cites.
Compressing language models using doped kronecker products
Thakker, U., Whatamough, P., Mattina, M., and Beu, J. G · 2001
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
An exploration of parameter redundancy in deep networks with circulant projections
Cheng, Y., Yu, F. X., Feris, R. S., Kumar, S., Choudhary, A., and Chang, S · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Structured transforms for small-footprint deep learning
Sindhwani, V., Sainath, T., and Kumar, S · 2015
Earlier work this paper cites.
Exploiting local structures with the kronecker layer in convolutional networks
Zhou, S., Wu, J., Wu, Y., and Zhou, X · 2015
Earlier work this paper cites.
Quasi-recurrent neural networks
Bradbury, J., Merity, S., Xiong, C., and Socher, R · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R · 2016
Earlier work this paper cites.
Attention-based recurrent neural network models for joint intent detection and slot filling
Liu, B. and Lane, I · 2016
Earlier work this paper cites.
Zilly, J. G., Srivastava, R. K., Koutník, J., and Schmidhuber, J · 2016
Earlier work this paper cites.
Neural networks compression for language modeling
Grachev, A. M., Ignatov, D. I., and Savchenko, A. V · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2017
Earlier work this paper cites.
Factorization tricks for LSTM networks
Kuchaiev, O. and Ginsburg, B · 2017
Earlier work this paper cites.
Neural machine translation (seq2seq) tutorial
Luong, M., Brevdo, E., and Zhao, R · 2017
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B · 2017
Earlier work this paper cites.
Structured bayesian pruning via log-normal multiplicative noise
Neklyudov, K., Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Cited alongside, same era.
Weighted-entropy-based quantization for deep neural networks
Park, E., Ahn, J., and Yoo, S · 2017
Cited alongside, same era.
Compressing recurrent neural network with tensor train
Tjandra, A., Sakti, S., and Nakamura, S · 2017
Cited alongside, same era.
Learning to skim text
Yu, A. W., Lee, H., and Le, Q · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
Skip rnn: Learning to skip state updates in recurrent neural networks
Campos, V., Jou, B., Giró-i Nieto, X., Torres, J., and Chang, S.-F · 2018
Ternary hybrid neural-tree networks for highly constrained iot applications
Gope, D., Dasika, G., and Mattina, M · 2019
Later among the works it cites.
Compression of recurrent neural networks for efficient language modeling
Grachev, A. M., Ignatov, D. I., and Savchenko, A. V · 2019
Later among the works it cites.
On-Chip Memory Technology Design Space Explorations for Mobile Deep Neural Network Accelerators
Li, H., Bhargav, M., Whatmough, P. N., and Philip Wong, H. · 2019
Later among the works it cites.
Define: Deep factorized input token embeddings for neural sequence modeling, 2019
Mehta, S., Koncel-Kedziorski, R., Rastegari, M., and Hajishirzi, H · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2019
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adaptive mixture of low-rank factorizations for compact neural modeling
Chen, T., Lin, J., Lin, T., Han, S., Wang, C., and Zhou, D · 2018
Cited alongside, same era.
Permdnn: Efficient compressed dnn architecture with permuted diagonal matrices
Deng, C., Liao, S., Xie, Y., Parhi, K. K., Qian, X., and Yuan, B · 2018
Cited alongside, same era.
Structured weight matrices-based hardware accelerators in deep neural networks: Fpgas and asics
Ding, C., Ren, A., Yuan, G., Ma, X., Li, J., Liu, N., Yuan, B., and Wang, Y · 2018
Cited alongside, same era.
Fastgrnn: A fast, accurate, stable and tiny kilobyte sized gated recurrent neural network
Kusupati, A., Singh, M., Bhatia, K., Kumar, A., Jain, P., and Varma, M · 2018
Cited alongside, same era.
Retraining-based iterative weight quantization for deep neural networks
Lee, D. and Kim, B · 2018
Cited alongside, same era.
Deeptwist: Learning model compression via occasional weight distortion
Lee, D., Kapoor, P., and Kim, B · 2018
Cited alongside, same era.
Tao, J., Thakker, U., Dasika, G., and Beu, J · 2019
Later among the works it cites.
Run-time efficient rnn compression for inference on edge devices
Thakker, U., Beu, J., Gope, D., Dasika, G., and Mattina, M · 2019
Later among the works it cites.
Structured pruning of recurrent neural networks through neuron selection
Wen, L., Zhang, X., Bai, H., and Xu, Z · 2019
Later among the works it cites.
FixyNN: Efficient Hardware for Mobile Computer Vision via Transfer Learning
Whatmough, P. N., Zhou, C., Hansen, P., Venkataramanaiah, S. K., sun Seo, J., and Mattina, M · 2019
Later among the works it cites.
Splitting steepest descent for growing neural architectures
Wu, L., Wang, D., and Liu, Q · 2019
Later among the works it cites.
Mobile Machine Learning Hardware at ARM: A Systems-on-Chip (SoC) Perspective
Zhu, Y., Mattina, M., and Whatmough, P. N · 2019
Later among the works it cites.
Exploring gpu acceleration of deep neural networks using block circulant matrices
Dong, S., Zhao, P., Lin, X., and Kaeli, D · 2020
Later among the works it cites.
Hawq-v2: Hessian aware trace-weighted quantization of neural networks
Dong, Z., Yao, Z., Arfeen, D., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
TinyLSTMs: Efficient Neural Speech Enhancement for Hearing Aids
Fedorov, I., Stamenovic, M., Jensen, C., Yang, L.-C., Mandell, A., Gan, Y., Mattina, M., and Whatmough, P. N · 2020
Later among the works it cites.
Raspberry pi 4 model b
Foundation, R. P · 2020
Later among the works it cites.
Eigen library
Guennebau, G. and Jacob, B · 2020
Later among the works it cites.
Pushing the envelope of dynamic spatial gating technologies
Huang, X., Thakker, U., Gope, D., and Beu, J · 2020
Later among the works it cites.
Systolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference
Liu, Z., Whatmough, P. N., and Mattina, M · 2020
Later among the works it cites.
Kronecker products
Nagy, J. G · 2020
Later among the works it cites.
Understanding the impact of dynamic channel pruning on conditionally parameterized convolutions
Raju, R., Gope, D., Thakker, U., and Beu, J · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning, 2020
Sanh, V., Wolf, T., and Rush, A. M · 2020
Later among the works it cites.
Rank and run-time aware compression of NLP applications
Thakker, U., Beu, J., Gope, D., Dasika, G., and Mattina, M · 2020
Later among the works it cites.
Micronets: Neural network architectures for deploying tinyml applications on commodity microcontrollers, 2021
Banbury, C., Zhou, C., Fedorov, I., Navarro, R. M., Thakker, U., Gope, D., Reddi, V. J., Mattina, M., and Whatmough, P. N · 2021
Closest in time.