Fetching the paper…
Reading the bibliography…
Today's highly heterogeneous computing landscape places a burden on programmers wanting to achieve high performance on a reasonably broad cross-section of machines.
Scans as primitive parallel operations
G. E. Blelloch · 1989
Earlier work this paper cites.
Automatic parallelization in the polytope model
P. Feautrier · 1996
Earlier work this paper cites.
Optimizing matrix multiply using PHiPAC: a portable, high-performance, ANSI C coding methodology
J. Bilmes, K. Asanovic, C.-W. Chin, and J. Demmel · 1997
Earlier work this paper cites.
Graphviz—open source graph drawing tools
J. Ellson, E. Gansner, L. Koutsofios, S. C. North, and G. Woodhull · 2002
Earlier work this paper cites.
Code generation in the polyhedral model is easier than you think
C. Bastoul · 2004
Earlier work this paper cites.
The landscape of parallel computing research: A view from berkeley
K. Asanovic, R. Bodik, B. C. Catanzaro, J. J. Gebis, P. Husbands, K. Keutzer, D. A. Patterson, W. L. Plishker, J. Shalf, S. W. Williams, et al · 2006
Earlier work this paper cites.
Loop transformation recipes for code generation and auto-tuning
M. Hall, J. Chame, C. Chen, J. Shin, G. Rudy, and M. Khan · 2010
Earlier work this paper cites.
hiCUDA: High-Level GPGPU Programming
T. D. Han and T. S. Abdelrahman · 2010
Earlier work this paper cites.
OpenMPC: Extended OpenMP programming and tuning for GPUs
S. Lee and R. Eigenmann · 2010
Cited alongside, same era.
GPGPU kernel implementation and refinement using Obsidian
J. Svensson, K. Claessen, and M. Sheeran · 2010
Cited alongside, same era.
isl: An integer set library for the polyhedral model
S. Verdoolaege · 2010
Cited alongside, same era.
A GPGPU compiler for memory optimization and parallelism management
Y. Yang, P. Xiang, J. Kong, and H. Zhou · 2010
Cited alongside, same era.
Copperhead: compiling an embedded data parallel language
B. Catanzaro, M. Garland, and K. Keutzer · 2011
Cited alongside, same era.
Automatic library generation for BLAS3 on GPUs
H. Cui, L. Wang, J. Xue, Y. Yang, and X. Feng · 2011
Cited alongside, same era.
The numpy array: a structure for efficient numerical computation
S. van der Walt, S. C. Colbert, and G. Varoquaux · 2011
Later among the works it cites.
A compiler toolkit for array-based languages targeting CPU/GPU hybrid systems
R. Garg and L. Hendren · 2012
Later among the works it cites.
Implementing a code generator for fast matrix multiplication in OpenCL on the GPU
K. Matsumoto, N. Nakasato, S. G. Sedukhin, I. M. Tsuruga, and A. W. City · 2012
Later among the works it cites.
Parakeet: A just-in-time parallel accelerator for Python
A. Rubinsteyn, E. Hielscher, N. Weinman, and D. Shasha · 2012
Later among the works it cites.
Polyhedral parallel code generation for CUDA
S. Verdoolaege, J. Carlos Juega, A. Cohen, J. Ignacio Gómez, C. Tenllado, and F. Catthoor · 2013
Later among the works it cites.
Numba Pro, 2014
Continuum Analytics, Inc · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PyCUDA and PyOpenCL: A Scripting-Based Approach to GPU Run-Time Code Generation
A. Klöckner, N. Pinto, Y. Lee, B. Catanzaro, P. Ivanov, and A. Fasih · 2011
Cited alongside, same era.
A programming language interface to describe transformations and code generation
G. Rudy, M. Khan, M. Hall, C. Chen, and J. Chame · 2011
Cited alongside, same era.
The islpy
A. Klöckner · 2014
Closest in time.
Loopy: Applications and Performance of transformation-based code generation for GPUs and CPUs
A. Klöckner and T. Warburton · 2014
Closest in time.