Fetching the paper…
Reading the bibliography…
We present automatic horizontal fusion, a novel optimization technique that complements the standard kernel fusion techniques for GPU programs.
J. Nickolls, I. Buck, M. Garland, and K. Skadron, “Scalable parallel programming with cuda,” Queue , vol. 6, no. 2, pp. 40–53, Mar. 2008. [Online]. Available: http://doi.acm.org/10.1145/1365490.1365500
2008
Earlier work this paper cites.
G. Wang, Y. Lin, and W. Yi, “Kernel fusion: An effective method for better power efficiency on multithreaded gpu,” in 2010 IEEE/ACM Int’l Conference on Green Computing and Communications Int’l Conference on Cyber, Physical and Social Computing , Dec 2010, pp. 344–350
2010
Earlier work this paper cites.
J. Fousek, J. Filipovič, and M. Madzin, “Automatic fusions of cuda-gpu kernels for parallel map,” SIGARCH Comput. Archit. News , vol. 39, no. 4, pp. 98–99, Dec. 2011. [Online]. Available: http://doi.acm.org/10.1145/2082156.2082183
2011
Earlier work this paper cites.
M. Bauer, H. Cook, and B. Khailany, “Cudadma: Optimizing gpu memory bandwidth via warp specialization,” in SC ’11: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis , 2011, pp. 1–11
2011
Earlier work this paper cites.
C. Gregg, J. Dorn, K. Hazelwood, and K. Skadron, “Fine-grained resource sharing for concurrent gpgpu kernels,” in Proceedings of the 4th USENIX Conference on Hot Topics in Parallelism , ser. HotPar’12. Berkeley, CA, USA: USENIX Association, 2012, pp. 10–10. [Online]. Available: http://dl.acm.org/citation.cfm?id=2342788.2342798
2012
Earlier work this paper cites.
S. Pai, M. J. Thazhuthaveetil, and R. Govindarajan, “Improving gpgpu concurrency with elastic kernels,” in Proceedings of the Eighteenth International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’13. New York, NY, USA: ACM, 2013, pp. 407–418. [Online]. Available: http://doi.acm.org/10.1145/2451116.2451160
2013
Earlier work this paper cites.
M. Wahib and N. Maruyama, “Scalable kernel fusion for memory-bound gpu applications,” in SC ’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , Nov 2014, pp. 191–202
2014
Earlier work this paper cites.
G. Wood, “Ethereum: A secure decentralised generalised transaction ledger,” Ethereum project yellow paper , vol. 151, pp. 1–32, 2014
2014
Earlier work this paper cites.
M. Bauer, S. Treichler, and A. Aiken, “Singe: Leveraging warp specialization for high performance on gpus,” in Proceedings of the 19th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , ser. PPoPP ’14. New York, NY, USA: ACM, 2014, pp. 119–130. [Online]. Available: http://doi.acm.org/10.1145/2555243.2555258
2014
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
J. Filipovič, M. Madzin, J. Fousek, and L. Matyska, “Optimizing cuda code by kernel fusion: application on blas,” The Journal of Supercomputing , vol. 71, no. 10, pp. 3934–3957, Oct 2015. [Online]. Available: https://doi.org/10.1007/s11227-015-1483-z
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 06 2016, pp. 770–778
2016
Cited alongside, same era.
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy, “TVM: An automated end-to-end optimizing compiler for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . Carlsbad, CA: USENIX Association, Oct. 2018, pp. 578–594. [Online]. Available: https://www.usenix.org/conference/osdi18/presentation/chen
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Wang, J. Yang, R. Melhem, B. Childers, Y. Zhang, and M. Guo, “Simultaneous multikernel gpu: Multi-tasking throughput processors via fine-grained sharing,” in 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA) , March 2016, pp. 358–369
2016
Cited alongside, same era.
2017
Cited alongside, same era.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in PyTorch,” in NIPS Autodiff Workshop , 2017
2017
Cited alongside, same era.
M. Springer, P. Wauligmann, and H. Masuhara, “Modular array-based gpu computing in a dynamically-typed language,” in Proceedings of the 4th ACM SIGPLAN International Workshop on Libraries, Languages, and Compilers for Array Programming , ser. ARRAY 2017. New York, NY, USA: ACM, 2017, pp. 48–55. [Online]. Available: http://doi.acm.org/10.1145/3091966.3091974
2017
Cited alongside, same era.
R. Ausavarungnirun, J. Landgraf, V. Miller, S. Ghose, J. Gandhi, C. J. Rossbach, and O. Mutlu, “Mosaic: A gpu memory manager with application-transparent support for multiple page sizes,” in 2017 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , Oct 2017, pp. 136–150
2017
Cited alongside, same era.
“Mlperf training v0.6 results,” https://mlperf.org/training-results-0-6/
Cited in the paper.
“Ethminer, ethereum miner with opencl, cuda and stratum support,” https://github.com/ethereum-mining/ethminer
Cited in the paper.
“ccminer, a cuda accelerated mining application,” https://github.com/tpruvot/ccminer
Cited in the paper.
R. Ausavarungnirun, V. Miller, J. Landgraf, S. Ghose, J. Gandhi, A. Jog, C. J. Rossbach, and O. Mutlu, “Mask: Redesigning the gpu memory hierarchy to support multi-application concurrency,” SIGPLAN Not. , vol. 53, no. 2, pp. 503–518, Mar. 2018. [Online]. Available: http://doi.acm.org/10.1145/3296957.3173169
2018
Later among the works it cites.
Z. Jia, O. Padon, J. Thomas, T. Warszawski, M. Zaharia, and A. Aiken, “Taso: Optimizing deep learning computation with automatic generation of graph substitutions,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , ser. SOSP ’19. New York, NY, USA: ACM, 2019, pp. 47–62. [Online]. Available: http://doi.acm.org/10.1145/3341301.3359630
2019
Later among the works it cites.
F. Boemer, Y. Lao, R. Cammarota, and C. Wierzynski, “ngraph-he: A graph compiler for deep learning on homomorphically encrypted data,” in Proceedings of the 16th ACM International Conference on Computing Frontiers , ser. CF ’19. New York, NY, USA: ACM, 2019, pp. 3–13. [Online]. Available: http://doi.acm.org/10.1145/3310273.3323047
2019
Later among the works it cites.
X. Li, S. Liu, S. D. Mello, X. Wang, J. Kautz, and M.-H. Yang, “Joint-task self-supervised learning for temporal correspondence,” in NeurIPS , 2019
2019
Later among the works it cites.
M. Sivathanu, T. Chugh, S. S. Singapuram, and L. Zhou, “Astra: Exploiting predictability to optimize deep learning,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’19. New York, NY, USA: ACM, 2019, pp. 909–923. [Online]. Available: http://doi.acm.org/10.1145/3297858.3304072
2019
Later among the works it cites.
G. Diamos, S. Sengupta, B. Catanzaro, M. Chrzanowski, A. Coates, E. Elsen, J. Engel, A. Hannun, and S. Satheesh, “Persistent rnns: Stashing recurrent weights on-chip,” in Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 , ser. ICML’16. JMLR.org, 2016, pp. 2024–2033. [Online]. Available: http://dl.acm.org/citation.cfm?id=3045390.3045604
2033
Closest in time.