Fetching the paper…
Reading the bibliography…
Despite significant investment in software infrastructure, machine learning systems, runtimes and compilers do not compose properly.
Theory of Linear and Integer Programming
Alexander Schrijver. 1986 · 1986
Earlier work this paper cites.
Semantical interprocedural parallelization: an overview of the PIPS project. In Proceedings of the 5th international conference on Supercomputing, ICS 1991, Cologne, Germany, June 17-21, 1991 , Edward S. Davidson and Friedel Hossfeld (Eds.). ACM, 244–251
François Irigoin, Pierre Jouvelot, and Rémi Triolet. 1991 · 1991
Earlier work this paper cites.
Some efficient solutions to the affine scheduling problem. Part I. One-dimensional time
Paul Feautrier. 1992a · 1992
Earlier work this paper cites.
Some efficient solutions to the affine scheduling problem. Part II. Multidimensional time
Paul Feautrier. 1992b · 1992
Earlier work this paper cites.
Combining Analyses, Combining Optimizations
Cliff Click and Keith D. Cooper. 1995 · 1995
Earlier work this paper cites.
Compiler Support for Sparse Matrix Computations
Aart J.C. Bik. 1996 · 1996
Earlier work this paper cites.
The Automatic Generation of Sparse Primitives
Aart J.C. Bik, Peter J.H. Brinkhaus, Peter M.W. Knijnenburg, and Harry A.G. Wijshoff. 1998 · 1998
Earlier work this paper cites.
Optimizing Compilers for Modern Architectures: A Dependence-Based Approach
Randy Allen and Ken Kennedy. 2001 · 2001
Earlier work this paper cites.
A Language for Describing Optimization Strategies
Bastian Hagedorn, Johannes Lenfers, Thomas Koehler, Sergei Gorlatch, and Michel Steuwer. 2020b · 2002
Earlier work this paper cites.
Fireiron: A Scheduling Language for High-Performance Linear Algebra on GPUs
Bastian Hagedorn, Archibald Samuel Elliott, Henrik Barthels, Rastislav Bodík, and Vinod Grover. 2020a · 2003
Earlier work this paper cites.
Code Generation in the Polyhedral Model Is Easier Than You Think. In 13th International Conference on Parallel Architectures and Compilation Techniques (PACT 2004), 29 September - 3 October 2004, Antibes Juan-les-Pins, France . IEEE Computer Society, 7–16
Cédric Bastoul. 2004 · 2004
Earlier work this paper cites.
Semi-Automatic Composition of Loop Transformations for Deep Parallelism and Memory Hierarchies
Sylvain Girbal, Nicolas Vasilache, Cédric Bastoul, Albert Cohen, David Parello, Marc Sigler, and Olivier Temam. 2006 · 2006
Earlier work this paper cites.
A practical automatic polyhedral parallelizer and locality optimizer. In Proceedings of the ACM SIGPLAN 2008 Conference on Programming Language Design and Implementation, Tucson, AZ, USA, June 7-13, 2008 , Rajiv Gupta and Saman P. Amarasinghe (Eds.). ACM, 101–113
Uday Bondhugula, Albert Hartono, J. Ramanujam, and P. Sadayappan. 2008 · 2008
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Samuel Williams, Andrew Waterman, and David A. Patterson. 2009 · 2009
Earlier work this paper cites.
Eigen v3
Gaël Guennebaud, Benoît Jacob, et al · 2010
Earlier work this paper cites.
GRAPHITE Two Years After: First Lessons Learned From Real-World Polyhedral Compilation. In GCC Research Opportunities Workshop (GROW’10) . HAL
Konrad Trifunovic, Albert Cohen, David Edelsohn, Feng Li, Tobias Grosser, Harsha Jagasia, Razya Ladelsky, Sebastian Pop, Jan Sjödin, and Ramakrishna Upadrasta. 2010 · 2010
Earlier work this paper cites.
isl : An Integer Set Library for the Polyhedral Model. In Mathematical Software – ICMS 2010, Third International Congress on Mathematical Software (Lecture Notes in Computer Science, Vol. 6327) , Komei Fukuda, Joris van der Hoeven, Michael Joswig, and Nobuki Takayama (Eds.). Springer, 299–302
Sven Verdoolaege. 2010 · 2010
Earlier work this paper cites.
R-Stream Compiler
Benoît Meister, Nicolas Vasilache, David Wohlford, Muthu Manikandan Baskaran, Allen Leung, and Richard Lethin. 2011 · 2011
Earlier work this paper cites.
Polly - Performing Polyhedral Optimizations on a Low-Level Intermediate Representation
Tobias Grosser, Armin Größlinger, and Christian Lengauer. 2012 · 2012
Earlier work this paper cites.
Joint scheduling and layout optimization to enable multi-level vectorization. In In Second International Workshop on Polyhedral Compilation Techniques (IMPACT) . Informal proceedings
Nicolas Vasilache, Benoit Meister, Muthu Baskaran, and Richard Lethin. 2012 · 2012
Earlier work this paper cites.
A script-based autotuning compiler system to generate high-performance CUDA code
Malik Murtaza Khan, Protonu Basu, Gabe Rudy, Mary W. Hall, Chun Chen, and Jacqueline Chame. 2013 · 2013
Cited alongside, same era.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Cited alongside, same era.
Practical SIMD Vectorization Techniques for Intel® Xeon Phi Coprocessors. In 2013 IEEE International Symposium on Parallel Distributed Processing, Workshops and Phd Forum . IEEE, 1149–1158
Xinmin Tian, Hideki Saito, Serguei V. Preis, Eric N. Garcia, Sergey S. Kozhukhov, Matt Masten, Aleksei G. Cherkasov, and Nikolay Panchenko. 2013 · 2013
Cited alongside, same era.
Caffe: Convolutional Architecture for Fast Feature Embedding. In Proceedings of the 22Nd ACM International Conference on Multimedia (Orlando, Florida, USA) (MM ’14) . ACM, 675–678
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. 2014 · 2014
Cited alongside, same era.
Machine Learning Systems Are Stuck in a Rut. In Proceedings of the Workshop on Hot Topics in Operating Systems (Bertinoro, Italy) (HotOS ’19) . Association for Computing Machinery, 177–183
Paul Barham and Michael Isard. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (Huntsville, Ontario, Canada) (SOSP ’19) . Association for Computing Machinery, 47–62
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken. 2019 · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PENCIL: A Platform-Neutral Compute Intermediate Language for Accelerator Programming. In 2015 International Conference on Parallel Architectures and Compilation, PACT 2015, San Francisco, CA, USA, October 18-21, 2015 . IEEE Computer Society, 138–149
Riyadh Baghdadi, Ulysse Beaugnon, Albert Cohen, Tobias Grosser, Michael Kruse, Chandan Reddy, Sven Verdoolaege, Adam Betts, Alastair F. Donaldson, Jeroen Ketema, Javed Absar, Sven van Haastregt, Alexey Kravets, Anton Lokhmotov, Robert David, and Elnar Hajiyev. 2015 · 2015
Cited alongside, same era.
Polyhedral AST Generation Is More Than Scanning Polyhedra
Tobias Grosser, Sven Verdoolaege, and Albert Cohen. 2015 · 2015
Cited alongside, same era.
Scientific Benchmarking of Parallel Computing Systems: Twelve Ways to Tell the Masses When Reporting Performance Results. In SC’15: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Austin, Texas). IEEE/ACM, Article 73, 12 pages
Torsten Hoefler and Roberto Belli. 2015 · 2015
Cited alongside, same era.
PolyMage: Automatic Optimization for Image Processing Pipelines. In Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS 2015, Istanbul, Turkey, March 14-18, 2015 , Özcan Özturk, Kemal Ebcioglu, and Sandhya Dwarkadas (Eds.). ACM, 429–443
Ravi Teja Mullapudi, Vinay Vasista, and Uday Bondhugula. 2015 · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
LIBXSMM: accelerating small matrix multiplications by runtime code generation. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE/ACM, 981–991
Alexander Heinecke, Greg Henry, Maxwell Hutchinson, and Hans Pabst. 2016 · 2016
Cited alongside, same era.
The Tensor Algebra Compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe. 2017 · 2017
Cited alongside, same era.
LIFT: A functional data-parallel IR for high-performance GPU code generation. In 2017 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE/ACM, 74–85
Michel Steuwer, Toomas Remmelg, and Christophe Dubach. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Swizzle Inventor: Data Movement Synthesis for GPU Kernels. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS 2019, Providence, RI, USA, April 13-17, 2019 , Iris Bahar, Maurice Herlihy, Emmett Witchel, and Alvin R. Lebeck (Eds.). ACM, 65–78
Phitchaya Mangpo Phothilimthana, Archibald Samuel Elliott, An Wang, Abhinav Jangda, Bastian Hagedorn, Henrik Barthels, Samuel J. Kaufman, Vinod Grover, Emina Torlak, and Rastislav Bodík. 2019 · 2019
Later among the works it cites.
Generating Portable High-Performance Code via Multi-Dimensional Homomorphisms. In 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 354–369
Ari Rasch, Richard Schulze, and Sergei Gorlatch. 2019 · 2019
Later among the works it cites.
The next 700 accelerated layers: From mathematical expressions of network computation graphs to accelerated GPU kernels, automatically
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary Devito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2019 · 2019
Later among the works it cites.
Achieving High-Performance the Functional Way: A Functional Pearl on Expressing High-Performance Optimizations as Rewrite Strategies
Bastian Hagedorn, Johannes Lenfers, Thomas Kundefinedhler, Xueying Qin, Sergei Gorlatch, and Michel Steuwer. 2020c · 2020
Later among the works it cites.
Ansor: Generating High-Performance Tensor Programs for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020, Virtual Event, November 4-6, 2020 . USENIX Association, 863–879
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica. 2020 · 2020
Later among the works it cites.
VeGen: a vectorizer generator for SIMD and beyond. In ASPLOS ’21: 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Virtual Event, USA, April 19-23, 2021 , Tim Sherwood, Emery D. Berger, and Christos Kozyrakis (Eds.). ACM, 902–914
Yishen Chen, Charith Mendis, Michael Carbin, and Saman P. Amarasinghe. 2021 · 2021
Later among the works it cites.
Intel Optimization Reference Manual
Intel Corp. 2021 · 2021
Later among the works it cites.
CompilerGym: Robust, Performant Compiler Optimization Environments for AI Research
Chris Cummins, Bram Wasti, Jiadong Guo, Brandon Cui, Jason Ansel, Sahir Gomez, Somya Jain, Jia Liu, Olivier Teytaud, Benoit Steiner, Yuandong Tian, and Hugh Leather. 2021 · 2021
Later among the works it cites.
IREE (Intermediate Representation Execution Environment
IREE Developers. 2021 · 2021
Later among the works it cites.
Benchmarking tips
LLVM Documentation. 2021 · 2021
Later among the works it cites.
Mlir: Scaling compiler infrastructure for domain specific computation. In 2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE/ACM, IEEE/ACM, 2–14
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko. 2021 · 2021
Later among the works it cites.
PDLL: a new declarative rewrite frontend for MLIR
River Riddle. 2021 · 2021
Later among the works it cites.
Outerproduct spilling
Giuseppe Rossini. 2021 · 2021
Later among the works it cites.
Pure Tensor Program Rewriting via Access Patterns (Representation Pearl)
Gus Henry Smith, Andrew Liu, Steven Lyubomirsky, Scott Davidson, Joseph McMahan, Michael B. Taylor, Luis Ceze, and Zachary Tatlock. 2021 · 2021
Later among the works it cites.
Understanding and controlling some of the AVX shuffle emission paths
Nicolas Vasilache. 2021 · 2021
Later among the works it cites.