Fetching the paper…
Reading the bibliography…
Performance optimization is the art of continuous seeking a harmonious mapping between the application domain and hardware.
Collective Loop Fusion for Array Contraction. In Proceedings of the 5th International Workshop on Languages and Compilers for Parallel Computing . London, UK
G. R. Gao, R. Olsen, V. Sarkar, and R. Thekkath. 1993 · 1993
Earlier work this paper cites.
Optimizing Compilers for Modern Architectures: A Dependence-based Approach
K. Kennedy and J. R. Allen. 2002 · 2002
Earlier work this paper cites.
“Improving Effective Bandwidth Through Compiler Enhancement of Global Cache Reuse
C. Ding and K. Kennedy. 2004 · 2004
Earlier work this paper cites.
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation. In Proceedings of 45th Annual IEEE/ACM International Symposium on Microarchitecture . Vancouver, BC, Canada
H. C. Wu, G. Diamos, S. Cadambi, and S. Yalamanchili. 2012 · 2012
Earlier work this paper cites.
Very Deep Convolutional Networks for Large Scale Image Recognition
K. Simonyan and A. Zisserman. 2014 · 2014
Earlier work this paper cites.
Scalable Kernel Fusion for Memory-Bound GPU Applications. In Proceedings of SC’14 . New Orleans, LA, USA
M. Wahib and N. Maruyama. 2014 · 2014
Earlier work this paper cites.
Torch NN
2015 · 2015
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2015 · 2015
Earlier work this paper cites.
Efficient Kernel Fusion Techniques for Massive Video Data Analysis on GPGPUs
A. M. Adnan, S. Radhakrishnan, and S. Karabuk. 2015 · 2015
Earlier work this paper cites.
On Optimizing Machine Learning Workloads via Kernel Fusion. In Proceedings of the 2015 ACM SIGPLAN Symposium on Principles and Practices of Parallel Programming . San Francisco, CA, USA
A. Ashari, S. Tatikonda, K. Campbell, and P. Sadayappan. 2015 · 2015
Cited alongside, same era.
MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015 · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
K. M. He, X. Y. Zhang, S. Q. Ren, and J. Sun. 2015 · 2015
Cited alongside, same era.
Rethinking the Inception Architecture for Computer Vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. 2015 · 2015
Cited alongside, same era.
Halide: decoupling algorithms from schedules for high-performance image processing
J. R. Kelley, A. Adams, D. Sharlet, C. Barnes, S. Paris, M. Levoy, S. Amarasinghe, and F. Durand. 2018 · 2018
Later among the works it cites.
Accelerating Explicit ODE Methods on GPUs by Kernel Fusion
M. Korch and T. Werner. 2018 · 2018
Later among the works it cites.
Automatic Kernel Fusion for Image Processing DSLs. In Proceedings of International Workshop on Software and Compilers for Embedded Systems . St. Goar, Germany
B. Qiao, O. Reiche, F. Hannig, and J. Teich. 2018 · 2018
Later among the works it cites.
Program Generation for Small-Scale Linear Algebra Applications. In Proceedings of the 2018 International Symposium on Code Generation and Optimization . Vienna, Austria
D. G. Spampinato, D. F. Traver, P. Bientinesi, and M. Püschel. 2018 · 2018
Later among the works it cites.
Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Latte: A Language, Compiler, and Runtime for Elegant and Efficient Deep Neural Networks. In Proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’16) . ACM, New York, NY, USA, 209–223
L. Truong, R. Barik, E. Totoni, H. Liu, C. Markley, A. Fox, and T. Shpeisman. 2016 · 2016
Cited alongside, same era.
Boda: A Holistic Approach for Implementing Neural Network Computations. In Proceedings of the Computing Frontiers Conference (CF’17) . ACM, New York, NY, USA, 53–62
M. W. Moskewicz, A. Jannesari, and K. Keutzer. 2017 · 2017
Cited alongside, same era.
Z. Zheng, C. Y. Oh, J. D. Zhai, X. P. Shen, and W. G. Chen. 2017 · 2017
Cited alongside, same era.
Optimal DNN Primitive Selection with Partitioned Boolean Quadratic Programming. In Proceedings of the 2018 International Symposium on Code Generation and Optimization . Vienna, Austria
A. Anderson and D. Gregg. 2018 · 2018
Cited alongside, same era.
TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of Operating Systems Design and Implemention (OSDI)
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy. 2018 · 2018
Cited alongside, same era.
TVM: Open Deep Learning Compiler Stack
[n. d.]
Cited in the paper.
XLA Operation Semantics
[n. d.]
Cited in the paper.
Tensorflow-Examples
aymericdamien. [n. d.]
Cited in the paper.
N. Vasilache, O. Zinenko, T. Theodoridis, P. Goyal, Z. DeVito, W. S. Moses, S. Verdoolaege, A. Adams, and A. Cohen. 2018 · 2018
Later among the works it cites.
Billion-scale Commodity Embedding for E-commerce Recommendation in Alibaba. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . London, United Kingdom
J. Z. Wang, P. P. Huang, H. Zhao, Z. B. Zhang, B. Q. Zhao, and D. L. Lee. 2018 · 2018
Later among the works it cites.
Attention Focusing for Neural Machine Translation by Bridging Source and Target Embeddings. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers . 1767–1776
D. Y. Xiong, J. H. Li, A. Branco, S. H. Kuang, and W. H. Luo. 2018 · 2018
Later among the works it cites.
From Loop Fusion to Kernel Fusion: A Domain Specific Approach to Locality Optimization. In Proceedings of International Symbopium on Code Generation and Optimization (CGO)
B. Qiao, O. Reiche, F. Hannig, and J. Teich. 2019 · 2019
Closest in time.
Astra: Exploiting Predictability to Optimize Deep Learning. In Proceedings of the Twenty Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) . 909–923
M. Slvathanu, T. Chugh, S. S. Singapuram, and L. D. Zhou. 2019 · 2019
Closest in time.