Fetching the paper…
Reading the bibliography…
The exponentially growing model size drives the continued success of deep learning, but it brings prohibitive computation and memory cost.
A. Pinar and M. T. Heath, “Improving performance of sparse matrix-vector multiplication,” in SC’99: Proceedings of the 1999 ACM/IEEE Conference on Supercomputing . IEEE, 1999, pp. 30–30
1999
Earlier work this paper cites.
R. Vuduc, J. W. Demmel, and K. A. Yelick, “Oski: A library of automatically tuned sparse matrix kernels,” in Journal of Physics: Conference Series , vol. 16, no. 1. IOP Publishing, 2005, p. 071
2005
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in European conference on computer vision . Springer, 2016, pp. 525–542
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” ACM SIGARCH Computer Architecture News , vol. 44, no. 3, pp. 243–254, 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
NVIDIA, “Nvidia tesla v100 gpu architecture,” 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
OpenAI, “AI and compute,” 2018. [Online]. Available: https://openai.com/blog/ai-and-compute/
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
F. Tung and G. Mori, “Deep neural network compression by in-parallel pruning-quantization,” IEEE transactions on pattern analysis and machine intelligence , vol. 42, no. 3, pp. 568–579, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Wang, W. Liu, W. Xue, and L. Wu, “swsptrsv: a fast sparse triangular solve with sparse level tile layout on sunway architectures,” in Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2018, pp. 338–353
2018
Earlier work this paper cites.
Y. Chen, K. Li, W. Yang, G. Xiao, X. Xie, and T. Li, “Performance-aware model for sparse matrix-matrix multiplication on the sunway taihulight supercomputer,” IEEE transactions on parallel and distributed systems , vol. 30, no. 4, pp. 923–938, 2018
2018
Earlier work this paper cites.
D. Zhang, J. Yang, D. Ye, and G. Hua, “Lq-nets: Learned quantization for highly accurate and compact deep neural networks,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 365–382
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in NLP,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL) , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Srivastava, D. Kadetotad, S. Yin, V. Berisha, C. Chakrabarti, and J.-s. Seo, “Joint optimization of quantization and structured sparsity for compressed deep neural networks,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 1393–1397
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat, “Q8bert: Quantized 8bit bert,” in 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS) . IEEE, 2019, pp. 36–39
2019
Earlier work this paper cites.
A. Li, T. Geng, T. Wang, M. Herbordt, S. L. Song, and K. Barker, “Bstc: A novel binarized-soft-tensor-core design for accelerating bit-based approximated neural nets,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2019, pp. 1–30
2019
Earlier work this paper cites.
Z. Xie, G. Tan, W. Liu, and N. Sun, “Ia-spgemm: An input-aware auto-tuning framework for parallel sparse matrix-matrix multiplication,” in Proceedings of the ACM International Conference on Supercomputing , 2019, pp. 94–105
2019
Cited alongside, same era.
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han, “Haq: Hardware-aware automated quantization with mixed precision,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8612–8620
2019
Cited alongside, same era.
2019
Cited alongside, same era.
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka, “Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 12 359–12 367
Z. Chen, Z. Qu, L. Liu, Y. Ding, and Y. Xie, “Efficient tensor core-based gpu kernels for structured sparsity under reduced precision,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–14
2021
Later among the works it cites.
NVIDIA, “Tuning CUDA Applications for Ampere,” September 2021, https://docs.nvidia.com/cuda/ampere-tuning-guide/index.html
2021
Later among the works it cites.
I. Hubara, Y. Nahshan, Y. Hanani, R. Banner, and D. Soudry, “Accurate post training quantization with small calibration sets,” in International Conference on Machine Learning . PMLR, 2021, pp. 4466–4475
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
——, “Exploiting NVIDIA Ampere Structured Sparsity with cuSPARSELt ,” December 2020, https://developer.nvidia.com/blog/exploiting-ampere-structured-sparsity-with-cusparselt/
2020
Cited alongside, same era.
T. Gale, M. Zaharia, C. Young, and E. Elsen, “Sparse gpu kernels for deep learning,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–14
2020
Cited alongside, same era.
M. Van Baalen, C. Louizos, M. Nagel, R. A. Amjad, Y. Wang, T. Blankevoort, and M. Welling, “Bayesian bits: Unifying quantization and pruning,” Advances in neural information processing systems , vol. 33, pp. 5741–5752, 2020
2020
Cited alongside, same era.
H. Yang, S. Gui, Y. Zhu, and J. Liu, “Automatic neural network compression by sparsity-quantization joint learning: A constrained optimization-based approach,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 2178–2188
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
J. Mao, H. Yang, A. Li, H. Li, and Y. Chen, “Tprune: Efficient transformer pruning for mobile devices,” ACM Transactions on Cyber-Physical Systems , vol. 5, no. 3, pp. 1–22, 2021
2021
Later among the works it cites.
S. Kim, A. Gholami, Z. Yao, M. W. Mahoney, and K. Keutzer, “I-bert: Integer-only bert quantization,” in International conference on machine learning . PMLR, 2021, pp. 5506–5518
2021
Later among the works it cites.
B. Feng, Y. Wang, T. Geng, A. Li, and Y. Ding, “Apnn-tc: Accelerating arbitrary precision neural networks on ampere gpu tensor cores,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–13
2021
Later among the works it cites.
C. Xie, J. Chen, J. Firoz, J. Li, S. L. Song, K. Barker, M. Raugas, and A. Li, “Fast and scalable sparse triangular solver for multi-gpu based hpc architectures,” in 50th International Conference on Parallel Processing , 2021, pp. 1–11
2021
Later among the works it cites.
Y. Niu, Z. Lu, M. Dong, Z. Jin, W. Liu, and G. Tan, “Tilespmv: A tiled algorithm for sparse matrix-vector multiplication on gpus,” in 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2021, pp. 68–78
2021
Later among the works it cites.
Z. Xie, G. Tan, W. Liu, and N. Sun, “A pattern-based spgemm library for multi-core and many-core architectures,” IEEE Transactions on Parallel and Distributed Systems , vol. 33, no. 1, pp. 159–175, 2021
2021
Later among the works it cites.
NVIDIA, “Parallel thread execution isa application guide,” 2021
2021
Later among the works it cites.
J. Liu, J. Ren, R. Gioiosa, D. Li, and J. Li, “Sparta: High-performance, element-wise sparse tensor contraction on heterogeneous memory,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 318–333
2021
Later among the works it cites.
NVIDIA, “CUDA C++ Programming Guide,” September 2021
2021
Later among the works it cites.
Google Research. [n.d], “Deep learning matrix collection,” 2021. [Online]. Available: https://github.com/google-research/google-research/tree/master/sgk
2021
Later among the works it cites.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 558–567
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “Vivit: A video vision transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6836–6846
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Li and T. Hoefler, “Chimera: efficiently training large-scale neural networks with bidirectional pipelines,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–14
2021
Later among the works it cites.
J. G. Pauloski, Q. Huang, L. Huang, S. Venkataraman, K. Chard, I. Foster, and Z. Zhang, “Kaisa: an adaptive second-order optimizer framework for deep neural networks,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–14
2021
Later among the works it cites.
Michael Andersch, Greg Palmer, Ronny Krashinsky, Nick Stam, Vishal Mehta, Gonzalo Brito and Sridhar Ramaswamy, “NVIDIA Hopper Architecture In-Depth,” March 2022. [Online]. Available: https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/
2022
Closest in time.
AMD, “AMD Instinct MI200 Instruction Set Architecture Reference Guide,” 2022
2022
Closest in time.
2022
Closest in time.
T. Wang, K. Wang, H. Cai, J. Lin, Z. Liu, H. Wang, Y. Lin, and S. Han, “Apq: Joint search for network architecture, pruning and quantization policy,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 2078–2087
2087
Closest in time.