Fetching the paper…
Reading the bibliography…
Achieving high-performance GPU kernels requires optimizing algorithm implementations to the targeted GPU architecture.
Building Program Optimizers with Rewriting Strategies. In Proceedings of the third ACM SIGPLAN International Conference on Functional Programming (ICFP ’98), Baltimore, Maryland, USA, September 27-29, 1998. 13–26
Eelco Visser, Zine-El-Abidine Benaissa, and Andrew P. Tolmach. 1998 · 1998
Earlier work this paper cites.
A Tactic Language for the System Coq. In Logic for Programming and Automated Reasoning, 7th International Conference, LPAR 2000, Reunion Island, France, November 11-12, 2000, Proceedings . 85–95
David Delahaye. 2000 · 2000
Earlier work this paper cites.
Rewriting with Strategies in ELAN: A Functional Semantics
Peter Borovanský, Claude Kirchner, Hélène Kirchner, and Christophe Ringeissen. 2001 · 2001
Earlier work this paper cites.
The ASF+SDF Meta-environment: A Component-Based Language Development Environment. In Compiler Construction, 10th International Conference, CC 2001 Held as Part of the Joint European Conferences on Theory and Practice of Software, ETAPS 2001 Genova, Italy, April 2-6, 2001, Proceedings . 365–370
Mark van den Brand, Arie van Deursen, Jan Heering, H. A. de Jong, Merijn de Jonge, Tobias Kuipers, Paul Klint, Leon Moonen, Pieter A. Olivier, Jeroen Scheerder, Jurgen J. Vinju, Eelco Visser, and Joost Visser. 2001 · 2001
Earlier work this paper cites.
A Language for the Compact Representation of Multiple Program Versions. In Languages and Compilers for Parallel Computing, 18th International Workshop, LCPC 2005, Hawthorne, NY, USA, October 20-22, 2005, Revised Selected Papers . 136–151
Sébastien Donadio, James C. Brodman, Thomas Roeder, Kamen Yotov, Denis Barthou, Albert Cohen, María Jesús Garzarán, David A. Padua, and Keshav Pingali. 2005 · 2005
Earlier work this paper cites.
CHiLL: A framework for composing high-level loop transformations
Chun Chen, Jacqueline Chame, and Mary Hall. 2008 · 2008
Earlier work this paper cites.
Annotation-based empirical performance tuning using Orio. In 23rd IEEE International Symposium on Parallel and Distributed Processing, IPDPS 2009, Rome, Italy, May 23-29, 2009 . 1–11
Albert Hartono, Boyana Norris, and Ponnuswamy Sadayappan. 2009 · 2009
Earlier work this paper cites.
Decoupling algorithms from schedules for easy optimization of image processing pipelines
Jonathan Ragan-Kelley, Andrew Adams, Sylvain Paris, Marc Levoy, Saman P. Amarasinghe, and Frédo Durand. 2012 · 2012
Earlier work this paper cites.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. In ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI ’13, Seattle, WA, USA, June 16-19, 2013 . 519–530
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman P. Amarasinghe. 2013 · 2013
Earlier work this paper cites.
Polyhedral parallel code generation for CUDA
Sven Verdoolaege, Juan Carlos Juega, Albert Cohen, José Ignacio Gómez, Christian Tenllado, and Francky Catthoor. 2013 · 2013
Cited alongside, same era.
OpenTuner: an extensible framework for program autotuning. In International Conference on Parallel Architectures and Compilation, PACT ’14, Edmonton, AB, Canada, August 24-27, 2014 . 303–316
Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman P. Amarasinghe. 2014 · 2014
Cited alongside, same era.
Schedule Trees. In Proceedings of the 4th International Workshop on Polyhedral Compilation Techniques , Sanjay Rajopadhye and Sven Verdoolaege (Eds.). Vienna, Austria
Sven Verdoolaege, Serge Guelton, Tobias Grosser, and Albert Cohen. 2014 · 2014
Cited alongside, same era.
Futhark: purely functional GPU-programming with nested parallelism and in-place array updates. In Proceedings of the 38th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2017, Barcelona, Spain, June 18-23, 2017 . 556–571
Troels Henriksen, Niels G. W. Serup, Martin Elsman, Fritz Henglein, and Cosmin E. Oancea. 2017 · 2017
Cited alongside, same era.
High performance stencil code generation with lift. In Proceedings of the 2018 International Symposium on Code Generation and Optimization, CGO 2018, Vösendorf / Vienna, Austria, February 24-28, 2018 . 100–112
Bastian Hagedorn, Larisa Stoltzfus, Michel Steuwer, Sergei Gorlatch, and Christophe Dubach. 2018 · 2018
Later among the works it cites.
A Proposal for Loop-Transformation Pragmas. In Evolving OpenMP for Evolving Architectures - 14th International Workshop on OpenMP, IWOMP 2018, Barcelona, Spain, September 26-28, 2018, Proceedings . 37–52
Michael Kruse and Hal Finkel. 2018 · 2018
Later among the works it cites.
Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Later among the works it cites.
GraphIt: a high-performance graph DSL
Yunming Zhang, Mengjiao Yang, Riyadh Baghdadi, Shoaib Kamil, Julian Shun, and Saman P. Amarasinghe. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ATF: A Generic Auto-Tuning Framework. In 19th IEEE International Conference on High Performance Computing and Communications; 15th IEEE International Conference on Smart City; 3rd IEEE International Conference on Data Science and Systems, HPCC/SmartCity/DSS 2017, Bangkok, Thailand, December 18-20, 2017 . 64–71
Ari Rasch, Michael Haidl, and Sergei Gorlatch. 2017 · 2017
Cited alongside, same era.
Lift: a functional data-parallel IR for high-performance GPU code generation. In Proceedings of the 2017 International Symposium on Code Generation and Optimization, CGO 2017, Austin, TX, USA, February 4-8, 2017 . 74–85
Michel Steuwer, Toomas Remmelg, and Christophe Dubach. 2017 · 2017
Cited alongside, same era.
TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2018, Carlsbad, CA, USA, October 8-10, 2018. 578–594
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018 · 2018
Cited alongside, same era.
Diesel: DSL for linear algebra and neural net computations on GPUs. In Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL@PLDI 2018, Philadelphia, PA, USA, June 18-22, 2018 . 42–51
Venmugil Elango, Norm Rubin, Mahesh Ravishankar, Hariharan Sandanagobalane, and Vinod Grover. 2018 · 2018
Cited alongside, same era.
Learning to optimize halide with tree search and random programs
Andrew Adams, Karima Ma, Luke Anderson, Riyadh Baghdadi, Tzu-Mao Li, Michaël Gharbi, Benoit Steiner, Steven Johnson, Kayvon Fatahalian, Frédo Durand, and Jonathan Ragan-Kelley. 2019 · 2019
Later among the works it cites.
Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable Code. In IEEE/ACM International Symposium on Code Generation and Optimization, CGO 2019, Washington, DC, USA, February 16-20, 2019 . 193–205
Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil, and Saman P. Amarasinghe. 2019 · 2019
Later among the works it cites.
Swizzle Inventor: Data Movement Synthesis for GPU Kernels. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS 2019, Providence, RI, USA, April 13-17, 2019 . 65–78
Phitchaya Mangpo Phothilimthana, Archibald Samuel Elliott, An Wang, Abhinav Jangda, Bastian Hagedorn, Henrik Barthels, Samuel J. Kaufman, Vinod Grover, Emina Torlak, and Rastislav Bodík. 2019 · 2019
Later among the works it cites.
Developing High-Performance, Portable OpenCL Code via Multi-Dimensional Homomorphisms. In Proceedings of the International Workshop on OpenCL, IWOCL 2019, Boston, MA, USA, May 13-15, 2019. 4:1
Ari Rasch, Richard Schulze, and Sergei Gorlatch. 2019 · 2019
Later among the works it cites.