Fetching the paper…
Reading the bibliography…
We present a new algorithm to quickly generate high-performance GPU implementations of complex imaging and vision pipelines, directly from high-level Halide algorithm code.
A practical automatic polyhedral parallelizer and locality optimizer
Uday Bondhugula, Albert Hartono, Jagannathan Ramanujam, and Ponnuswamy Sadayappan. 2008 · 2008
Earlier work this paper cites.
Polly — performing polyhedral optimizations on a low-level intermediate representation
Tobias Grosser, Armin Groesslinger, and Christian Lengauer. 2012 · 2012
Earlier work this paper cites.
Decoupling Algorithms from Schedules for Easy Optimization of Image Processing Pipelines
Jonathan Ragan-Kelley, Andrew Adams, Sylvain Paris, Marc Levoy, Saman Amarasinghe, and Frédo Durand. 2012 · 2012
Earlier work this paper cites.
Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Earlier work this paper cites.
Polyhedral Parallel Code Generation for CUDA
Sven Verdoolaege, Juan Carlos Juega, Albert Cohen, José Ignacio Gómez, Christian Tenllado, and Francky Catthoor. 2013 · 2013
Earlier work this paper cites.
OpenTuner: An extensible framework for program autotuning. In Parallel Architectures and Compilation . ACM, 303–316
Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman Amarasinghe. 2014 · 2014
Earlier work this paper cites.
Learning visual representations at scale
Vincent Vanhoucke. 2014 · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2015 · 2015
Earlier work this paper cites.
PolyMage: Automatic Optimization for Image Processing Pipelines
Ravi Teja Mullapudi, Vinay Vasista, and Uday Bondhugula. 2015 · 2015
Earlier work this paper cites.
Automatically Scheduling Halide Image Processing Pipelines
Ravi Teja Mullapudi, Andrew Adams, Dillon Sharlet, Jonathan Ragan-Kelley, and Kayvon Fatahalian. 2016 · 2016
Cited alongside, same era.
A survey on compiler autotuning using machine learning
Amir H Ashouri, William Killian, John Cavazos, Gianluca Palermo, and Cristina Silvano. 2018 · 2018
Cited alongside, same era.
An Effective Fusion and Tile Size Model for Optimizing Image Processing Pipelines
Abhinav Jangda and Uday Bondhugula. 2018 · 2018
Cited alongside, same era.
Differentiable programming for image processing and deep learning in Halide
Tzu-Mao Li, Michaël Gharbi, Andrew Adams, Frédo Durand, and Jonathan Ragan-Kelley. 2018 · 2018
Cited alongside, same era.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Cited alongside, same era.
Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation. In International Conference on Learning Representations
Byung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, and Hadi Esmaeilzadeh. 2020 · 2020
Closest in time.
NeuroVectorizer: end-to-end vectorization with deep reinforcement learning. 242–255
Ameer Haj-Ali, Nesreen K Ahmed, Ted Willke, Yakun Sophia Shao, Krste Asanovic, and Ion Stoica. 2020 · 2020
Closest in time.
Reinforced Genetic Algorithm Learning for Optimizing Computation Graphs. In Proceedings of the International Conference on Learning Representations (ICLR)
Aditya Paliwal, Felix Gimeno, Vinod Nair, Yujia Li, Miles Lubin, Pushmeet Kohli, and Oriol Vinyals. 2020 · 2020
Closest in time.
Schedule Synthesis for Halide Pipelines on GPUs
Savvas Sioutas, Sander Stuijk, Twan Basten, Henk Corporaal, and Lou Somers. 2020 · 2020
Closest in time.
XLA – TensorFlow compiled
The XLA Team. 2017 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to Optimize Halide with Tree Search and Random Programs
Andrew Adams, Karima Ma, Luke Anderson, Riyadh Baghdadi, Tzu-Mao Li, Michaël Gharbi, Benoit Steiner, Steven Johnson, Kayvon Fatahalian, Frédo Durand, and Jonathan Ragan-Kelley. 2019 · 2019
Cited alongside, same era.
TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions. In Proceedings of the ACM Symposium on Operating Systems Principles (SOSP) . ACM, 47–62
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken. 2019 · 2019
Cited alongside, same era.
Ithemal: Accurate, portable and fast basic block throughput estimation using deep neural networks. In International Conference on Machine Learning . PMLR, 4505–4515
Charith Mendis, Alex Renda, Saman Amarasinghe, and Michael Carbin. 2019 · 2019
Cited alongside, same era.
Schedule Synthesis for Halide Pipelines through Reuse Analysis
Savvas Sioutas, Sander Stuijk, Luc Waeijen, Twan Basten, Henk Corporaal, and Lou Somers. 2019 · 2019
Cited alongside, same era.
TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (Carlsbad, CA, USA) (OSDI’18) . USENIX Association, USA, 579–594
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018a
Cited in the paper.
Learning to optimize tensor programs. In Advances in Neural Information Processing Systems . 3389–3400
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018b
Cited in the paper.
FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System. In International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) . 859–873
Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng. 2020b
Cited in the paper.
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et al · 2020
Closest in time.
Transferable Graph Optimizers for ML Compilers. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. 13844–13855
Yanqi Zhou, Sudip Roy, Amirali Abdolrashidi, Daniel Wong, Peter Ma, Qiumin Xu, Hanxiao Liu, Phitchaya Phothilimtha, Shen Wang, Anna Goldie, Azalia Mirhoseini, and James Laudon. 2020 · 2020
Closest in time.
A Deep Learning Based Cost Model for Automatic Code Optimization. In Proceedings of the Fourth Conference on Machine Learning and Systems (San Jose, CA, USA) (MLSys 2021)
Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham Leghettas, Kamel Abdous, Taha Arbaoui, Karima Benatchba, and Saman Amarasinghe. 2021 · 2021
Closest in time.
Value Learning for Throughput Optimization of Deep Learning Workloads. In Proceedings of Machine Learning and Systems
Benoit Steiner, Chris Cummins, Horace He, and Hugh Leather. 2021 · 2021
Closest in time.