Fetching the paper…
Reading the bibliography…
Scientific workloads have traditionally exploited high levels of sparsity to accelerate computation and reduce memory requirements.
C. D. Polychronopoulos and D. J. Kuck, “Guided self-scheduling: A practical scheduling scheme for parallel supercomputers,” IEEE Trans. Computers , vol. 36, no. 12, pp. 1425–1439, 1987
1987
Earlier work this paper cites.
Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal Brain Damage,” in Advances in Neural Information Processing Systems 2, [NIPS Conference, Denver, Colorado, USA, November 27-30, 1989] , 1989
1989
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Buluç, J. T. Fineman, M. Frigo, J. R. Gilbert, and C. E. Leiserson, “Parallel sparse matrix-vector and matrix-transpose-vector multiplication using compressed sparse blocks,” in SPAA 2009: Proceedings of the 21st Annual ACM Symposium on Parallelism in Algorithms and Architectures, Calgary, Alberta, Canada, August 11-13, 2009 , F. M. auf der Heide and M. A. Bender, Eds. ACM, 2009, pp. 233–244
2009
Earlier work this paper cites.
M. Naumov, L. S. Chien, P. Vandermersch, and U. Kapasi, “cuSPARSE Library,” https://www.nvidia.com/content/GTC-2010/pdfs/2070_GTC2010.pdf , 2010
2010
Earlier work this paper cites.
T. A. Davis and Y. Hu, “The university of florida sparse matrix collection,” ACM Trans. Math. Softw. , vol. 38, no. 1, 2011
2011
Earlier work this paper cites.
F. Vázquez, J. Fernández, and E. M. Garzón, “A new approach for sparse matrix vector product on NVIDIA gpus,” Concurr. Comput. Pract. Exp. , vol. 23, no. 8, pp. 815–826, 2011
2011
Earlier work this paper cites.
P. Micikevicius, “GPU Performance Analysis and Optimization,” http://on-demand.gputechconf.com/gtc/2012/presentations/S0514-GTC2012-GPU-Performance-Analysis.pdf , 2012
2012
Earlier work this paper cites.
F. Vázquez, G. O. López, J. Fernández, I. García, and E. M. Garzón, “Fast sparse matrix matrix product based on ELLR-T and GPU computing,” in 10th IEEE International Symposium on Parallel and Distributed Processing with Applications, ISPA 2012, Leganes, Madrid, Spain, July 10-13, 2012 . IEEE Computer Society, 2012, pp. 669–674
2012
Earlier work this paper cites.
J. Luitjens, “CUDA Pro Tip: Increase Performance with Vectorized Memory Access,” https://devblogs.nvidia.com/cuda-pro-tip-increase-performance-with-vectorized-memory-access/ , 2013
2013
Earlier work this paper cites.
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. Tamchyna, “Findings of the 2014 workshop on statistical machine translation,” in Proceedings of the Ninth Workshop on Statistical Machine Translation . Baltimore, Maryland, USA: Association for Computational Linguistics, June 2014, pp. 12–58. [Online]. Available: http://www.aclweb.org/anthology/W/W14/W14-3302
2014
Earlier work this paper cites.
S. Pai, “How the Fermi Thread Block Scheduler Works,” https://www.cs.rochester.edu/~sree/fermi-tbs/fermi-tbs.html , 2014
2014
Earlier work this paper cites.
A. Adinets, “Adaptive parallel computation with cuda dynamic parallelism,” https://devblogs.nvidia.com/introduction-cuda-dynamic-parallelism/ , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL , A. Moschitti, B. Pang, and W. Daelemans, Eds. ACL, 2014, pp. 1724–1734
2014
Earlier work this paper cites.
H. Anzt, S. Tomov, and J. J. Dongarra, “Implementing a sparse matrix vector product for the sell-c / sell-c- σ \sigma formats on nvidia gpus,” 2014
2014
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both Weights and Connections for Efficient neural network,” in Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li, “Imagenet large scale visual recognition challenge,” Int. J. Comput. Vis. , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
H. Anzt, S. Tomov, and J. J. Dongarra, “Accelerating the LOBPCG method on gpus using a blocked sparse matrix vector product,” in Proceedings of the Symposium on High Performance Computing, HPC 2015, part of the 2015 Spring Simulation Multiconference, SpringSim ’15, Alexandria, VA, USA, April 12-15, 2015 . SCS/ACM, 2015, pp. 75–82
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 , ser. JMLR Workshop and Conference Proceedings, F. R. Bach and D. M. Blei, Eds., vol. 37. JMLR.org, 2015, pp. 448–456
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 , 2016
2016
Cited alongside, same era.
S. Narang, “DeepBench: Benchmarking Deep Learning Operations on Different Hardware,” https://github.com/baidu-research/DeepBench , 2016
2016
Cited alongside, same era.
D. Merrill and M. Garland, “Merge-Based Parallel Sparse Matrix-Vector Multiplication,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2016, Salt Lake City, UT, USA, November 13-18, 2016 , 2016
C. Yang, A. Buluç, and J. D. Owens, “Design Principles for Sparse Matrix Multiplication on the GPU,” in Euro-Par 2018: Parallel Processing - 24th International Conference on Parallel and Distributed Computing, Turin, Italy, August 27-31, 2018, Proceedings , 2018, pp. 672–687
2018
Later among the works it cites.
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image Transformer,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 , 2018
2018
Later among the works it cites.
F. Zhu, J. Pool, M. Andersch, J. Appleyard, and F. Xie, “Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
A. Lavin and S. Gray, “Fast algorithms for convolutional neural networks,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 . IEEE Computer Society, 2016, pp. 4013–4021
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA , 2017
2017
Cited alongside, same era.
S. Narang, G. Diamos, S. Sengupta, and E. Elsen, “Exploring Sparsity in Recurrent Neural Networks,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings , 2017
2017
Cited alongside, same era.
D. Molchanov, A. Ashukha, and D. P. Vetrov, “Variational dropout sparsifies deep neural networks,” in Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Gray, A. Radford, and D. P. Kingma, “Block-sparse gpu kernels,” https://blog.openai.com/block-sparse-gpu-kernels/ , 2017
2017
Cited alongside, same era.
A. Kerr, D. Merrill, J. Demouth, and J. Tran, “CUTLASS: Fast Linear Algebra in CUDA C++,” https://devblogs.nvidia.com/cutlass-linear-algebra-cuda/ , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, D. García, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings , 2018
2018
Later among the works it cites.
M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, and L. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018 . IEEE Computer Society, 2018, pp. 4510–4520
2018
Later among the works it cites.
J. Li, J. Sun, and R. W. Vuduc, “Hicoo: hierarchical storage of sparse tensors,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2018, Dallas, TX, USA, November 11-16, 2018 . IEEE / ACM, 2018, pp. 19:1–19:15
2018
Later among the works it cites.
2019
Later among the works it cites.
Z. Yao, S. Cao, W. Xiao, C. Zhang, and L. Nie, “Balanced sparsity for efficient DNN inference on GPU,” in The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019 , 2019, pp. 5676–5683
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Hong, A. Sukumaran-Rajam, I. Nisa, K. Singh, and P. Sadayappan, “Adaptive sparse tiling for sparse matrix multiplication,” in Proceedings of the 24th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, PPoPP 2019, Washington, DC, USA, February 16-20, 2019 , J. K. Hollingsworth and I. Keidar, Eds. ACM, 2019, pp. 300–314
2019
Later among the works it cites.
M. Tan and Q. V. Le, “Efficientnet: Rethinking Model Scaling for Convolutional Neural Networks,” in Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA , 2019
2019
Later among the works it cites.
Nvidia, “CUDA C++ Programming Guide,” https://docs.nvidia.com/cuda/cuda-c-programming-guide/ , 2019
2019
Later among the works it cites.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , A. Korhonen, D. R. Traum, and L. Màrquez, Eds. Association for Computational Linguistics, 2019, pp. 2978–2988
2019
Later among the works it cites.
A. Howard, R. Pang, H. Adam, Q. V. Le, M. Sandler, B. Chen, W. Wang, L. Chen, M. Tan, G. Chu, V. Vasudevan, and Y. Zhu, “Searching for mobilenetv3,” in 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 . IEEE, 2019, pp. 1314–1324
2019
Later among the works it cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The Efficient Transformer,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=rkgNKkHtvB
2020
Closest in time.
2020
Closest in time.
Nvidia, “NVIDIA A100 tensor core gpu architecture,” https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf , 2020
2020
Closest in time.