Fetching the paper…
Reading the bibliography…
A commonly occurring computation idiom in neural networks is to perform some pointwise operations on the result of a matrix multiplication.
1903
Earlier work this paper cites.
P. Feautrier, “Some efficient solutions to the affine scheduling problem: Part I, one-dimensional time,” Intl. Journal of Parallel Programming , vol. 21, no. 5, pp. 313–348, 1992
1992
Earlier work this paper cites.
——, “Some efficient solutions to the affine scheduling problem: Part II, multidimensional time,” Intl. Journal of Parallel Programming , vol. 21, no. 6, pp. 389–420, 1992
1992
Earlier work this paper cites.
2002
Earlier work this paper cites.
S. Pop, A. Cohen, C. Bastoul, S. Girbal, G. Silber, and N. Vasilache, “Graphite: Loop optimizations based on the polyhedral model for gcc,” 2006
2006
Earlier work this paper cites.
U. Bondhugula, A. Hartono, J. Ramanujam, and P. Sadayappan, “A practical automatic polyhedral program optimization system,” in PLDI , Jun 2008
2008
Earlier work this paper cites.
M. Baskaran, U. Bondhugula, S. Krishnamoorthy, J. Ramanujam, A. Rountev, and P. Sadayappan, “A Compiler Framework for Optimization of Affine Loop Nests for GPGPUs,” in ACM Intl. conference on Supercomputing (ICS) , Jun. 2008
2008
Earlier work this paper cites.
S. Verdoolaege, “ isl : An integer set library for the polyhedral model,” in Mathematical Software - ICMS 2010, Third International Congress on Mathematical Software, Kobe, Japan, September 13-17, 2010. Proceedings , 2010, pp. 299–302. [Online]. Available: https://doi.org/10.1007/978-3-642-15582-6_49
2010
Earlier work this paper cites.
T. Grosser, H. Zheng, R. Aloor, A. Simbürger, A. Großlinger, and L.-N. Pouchet, “Polly: Polyhedral optimization in LLVM,” in IMPACT , 2011
2011
Earlier work this paper cites.
S. Verdoolaege, J. C. Juega, A. Cohen, J. I. Gómez, C. Tenllado, and F. Catthoor, “Polyhedral parallel code generation for CUDA,” TACO , vol. 9, no. 4, pp. 54:1–54:23, 2013. [Online]. Available: https://doi.org/10.1145/2400682.2400713
2013
Earlier work this paper cites.
J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, and S. P. Amarasinghe, “Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines,” in ACM SIGPLAN symposium on Programming Languages Design and Implementation , 2013, pp. 519–530
2013
Cited alongside, same era.
2014
Cited alongside, same era.
S. Verdoolaege, S. Guelton, T. Grosser, and A. Cohen, “Schedule trees,” in IMPACT , 01 2014
2014
Cited alongside, same era.
R. T. Mullapudi, V. Vasista, and U. Bondhugula, “Polymage: Automatic optimization for image processing pipelines,” in Intl. Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’15, 2015, pp. 429–443
2015
Cited alongside, same era.
2018
Later among the works it cites.
NVIDIA, “cublas,” 2019, https://docs.nvidia.com/cuda/cublas/index.html
2019
Later among the works it cites.
——, “Cuda templates for linear algebra subroutines,” 2019, https://github.com/NVIDIA/cutlass
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 December 2019, Vancouver, BC, Canada , 2019, pp. 8024–8035. [Online]. Available: http://papers.nips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. A. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng, “Tensorflow: A system for large-scale machine learning,” in 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016, Savannah, GA, USA, November 2-4, 2016 , 2016, pp. 265–283. [Online]. Available: https://www.usenix.org/conference/osdi16/technical-sessions/presentation/abadi
2016
Cited alongside, same era.
——, “Programming tensor cores in cuda 9,” 2017, https://devblogs.nvidia.com/programming-tensor-cores-cuda-9/
2017
Cited alongside, same era.
2018
Cited alongside, same era.
V. Elango, N. Rubin, M. Ravishankar, H. Sandanagobalane, and V. Grover, “Diesel: DSL for linear algebra and neural net computations on gpus,” in Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL@PLDI 2018, Philadelphia, PA, USA, June 18-22, 2018 , 2018, pp. 42–51. [Online]. Available: https://doi.org/10.1145/3211346.3211354
2018
Cited alongside, same era.
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Q. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy, “TVM: an automated end-to-end optimizing compiler for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2018, Carlsbad, CA, USA, October 8-10, 2018 , 2018, pp. 578–594. [Online]. Available: https://www.usenix.org/conference/osdi18/presentation/chen
2018
Cited alongside, same era.
“How to optimize convolution using tensorcores,” 2018, https://docs.tvm.ai/tutorials/optimize/opt_conv_tensorcore.html
2018
Cited alongside, same era.
“PLUTO: An automatic polyhedral parallelizer and locality optimizer for multicores,” http://pluto-compiler.sourceforge.net
Cited in the paper.
“POCC: Polyhedral compiler collection,” http://pocc.sourceforge.net
Cited in the paper.
2019
Later among the works it cites.
NVIDIA, “Cuda toolkit documentation,” 2019, https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-fragment-mma-884
2019
Later among the works it cites.
NVIDIA, “cutensor: A high-performance cuda library for tensor primitives,” 2019, https://docs.nvidia.com/cuda/cutensor/index.html
2019
Later among the works it cites.
R. Baghdadi, J. Ray, M. B. Romdhane, E. D. Sozzo, A. Akkas, Y. Zhang, P. Suriana, S. Kamil, and S. P. Amarasinghe, “Tiramisu: A polyhedral compiler for expressing fast and portable code,” in IEEE/ACM International Symposium on Code Generation and Optimization, CGO 2019, Washington, DC, USA, February 16-20, 2019 , 2019, pp. 193–205. [Online]. Available: https://doi.org/10.1109/CGO.2019.8661197
2019
Later among the works it cites.
Intel, “Plaidml,” 2019, https://www.intel.ai/plaidml
2019
Later among the works it cites.
S. V. M. K. R. Schreiber and H. Kamepalli, “Generating simd instructions for cerebras cs-1 using polyhedral compilation techniques,” 2020
2020
Closest in time.
N. Vasilache, O. Zinenko, T. Theodoridis, P. Goyal, Z. DeVito, W. S. Moses, S. Verdoolaege, A. Adams, and A. Cohen, “The next 700 accelerated layers: From mathematical expressions of network computation graphs to accelerated GPU kernels, automatically,” TACO , vol. 16, no. 4, pp. 38:1–38:26, 2020. [Online]. Available: https://doi.org/10.1145/3355606
2020
Closest in time.