Fetching the paper…
Reading the bibliography…
We show in this work that memory intensive computations can result in severe performance problems due to off-chip memory access and CPU-GPU context switch overheads in a wide range of deep learning models.
Collective loop fusion for array contraction
G Gao, Russ Olsen, Vivek Sarkar, and Radhika Thekkath · 1992
Earlier work this paper cites.
A simple, fast dominance algorithm, 2001
Keith D Cooper, Timothy J Harvey, and Ken Kennedy · 2001
Earlier work this paper cites.
Optimizing compilers for modern architectures: a dependence-based approach
Randy Allen and Ken Kennedy · 2002
Earlier work this paper cites.
Limited discrepancy beam search
David Furcy and Sven Koenig · 2005
Earlier work this paper cites.
A quantitative performance analysis model for gpu architectures
Yao Zhang and John D Owens · 2011
Earlier work this paper cites.
An accurate gpu performance model for effective control flow divergence optimization
Zheng Cui, Yun Liang, Kyle Rupnow, and Deming Chen · 2012
Earlier work this paper cites.
Kernel weaver: Automatically fusing database primitives for efficient gpu computation
Haicheng Wu, Gregory Diamos, Srihari Cadambi, and Sudhakar Yalamanchili · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Scalable kernel fusion for memory-bound gpu applications
Mohamed Wahib and Naoya Maruyama · 2014
Earlier work this paper cites.
Torch nn, 2015
2015
Earlier work this paper cites.
Efficient kernel fusion techniques for massive video data analysis on gpgpus, 2015
Asif M Adnan, Sridhar Radhakrishnan, and Suleyman Karabuk · 2015
Earlier work this paper cites.
On optimizing machine learning workloads via kernel fusion
Arash Ashari, Shirish Tatikonda, Matthias Boehm, Berthold Reinwald, Keith Campbell, John Keenleyside, and P Sadayappan · 2015
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems, 2016
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Earlier work this paper cites.
Optimizing performance of recurrent neural networks on gpus
Jeremy Appleyard, Tomas Kocisky, and Phil Blunsom · 2016
Earlier work this paper cites.
Persistent rnns: Stashing recurrent weights on-chip
Greg Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Latte: a language, compiler, and runtime for elegant and efficient deep neural networks
Leonard Truong, Rajkishore Barik, Ehsan Totoni, Hai Liu, Chick Markley, Armando Fox, and Tatiana Shpeisman · 2016
Cited alongside, same era.
AUTOMATIC SPEECH RECOGNITION
Dong Yu and Li Deng · 2016
Cited alongside, same era.
Xla: Tensorflow, compiled, 2017
Chris Leary and Todd Wang · 2017
Cited alongside, same era.
Boda: A holistic approach for implementing neural network computations
Matthew W Moskewicz, Ali Jannesari, and Kurt Keutzer · 2017
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions, 2018
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen · 2018
Later among the works it cites.
Billion-scale commodity embedding for e-commerce recommendation in alibaba
Jizhe Wang, Pipei Huang, Huan Zhao, Zhibo Zhang, Binqiang Zhao, and Dik Lun Lee · 2018
Later among the works it cites.
Learning to fuse
Amirali Abdolrashidi, Qiumin Xu, Shibo Wang, Sudip Roy, and Yanqi Zhou · 2019
Later among the works it cites.
Getting started with cuda graphs
A Gray · 2019
Later among the works it cites.
Dissecting the nvidia turing t4 gpu via microbenchmarking
Zhe Jia, Marco Maggioni, Jeffrey Smith, and Daniele Paolo Scarpazza · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Halide: decoupling algorithms from schedules for high-performance image processing
Jonathan Ragan-Kelley, Andrew Adams, Dillon Sharlet, Connelly Barnes, Sylvain Paris, Marc Levoy, Saman Amarasinghe, and Frédo Durand · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Understanding the gpu microarchitecture to achieve bare-metal performance tuning
Xiuxia Zhang, Guangming Tan, Shuangbai Xue, Jiajia Li, Keren Zhou, and Mingyu Chen · 2017
Cited alongside, same era.
Versapipe: a versatile programming framework for pipelined computing on gpu
Zhen Zheng, Chanyoung Oh, Jidong Zhai, Xipeng Shen, Youngmin Yi, and Wenguang Chen · 2017
Cited alongside, same era.
Optimal dnn primitive selection with partitioned boolean quadratic programming
Andrew Anderson and David Gregg · 2018
Cited alongside, same era.
Jinsung Kim, Aravind Sukumaran-Rajam, Vineeth Thumma, Sriram Krishnamoorthy, Ajay Panyala, Louis-Noël Pouchet, Atanas Rountev, and Ponnuswamy Sadayappan · 2019
Later among the works it cites.
Delta: Gpu performance model for deep learning applications with in-depth memory system traffic analysis
Sangkug Lym, Donghyuk Lee, Mike O’Connor, Niladrish Chatterjee, and Mattan Erez · 2019
Later among the works it cites.
From loop fusion to kernel fusion: a domain-specific approach to locality optimization
Bo Qiao, Oliver Reiche, Frank Hannig, and Jïrgen Teich · 2019
Later among the works it cites.
Astra: Exploiting predictability to optimize deep learning
Muthian Sivathanu, Tapan Chugh, Sanjay S Singapuram, and Lidong Zhou · 2019
Later among the works it cites.
Astra: Exploiting predictability to optimize deep learning
Muthian Sivathanu, Tapan Chugh, Sanjay S Singapuram, and Lidong Zhou · 2019
Later among the works it cites.
A performance model for gpu architectures that considers on-chip resources: application to medical image registration
Xuan Yang, Zhengrui Zhang, Guoliang Chen, Rui Mao, et al · 2019
Later among the works it cites.
Deep interest evolution network for click-through rate prediction
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai · 2019
Later among the works it cites.
Automatic generation of high-performance quantized machine learning kernels
Meghan Cowan, Thierry Moreau, Tianqi Chen, James Bornholt, and Luis Ceze · 2020
Closest in time.
A learned performance model for the tensor processing unit
Samuel J Kaufman, Phitchaya Mangpo Phothilimthana, Yanqi Zhou, and Mike Burrows · 2020
Closest in time.
Automatic horizontal fusion for gpu kernels
Ao Li, Bojian Zheng, Gennady Pekhimenko, and Fan Long · 2020
Closest in time.
Dag-based scheduling with resource sharing for multi-task applications in a polyglot gpu runtime
Alberto Parravicini, Arnaud Delamare, Marco Arnaboldi, and Marco D Santambrogio · 2020
Closest in time.
Ansor: Generating high-performance tensor programs for deep learning
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et al · 2020
Closest in time.