Fetching the paper…
Reading the bibliography…
High-performance computing has recently seen a surge of interest in heterogeneous systems, with an emphasis on modern Graphics Processing Units (GPUs).
Methods of conjugate gradients for solving linear systems
M. R. Hestenes and E. Stiefel · 1952
Earlier work this paper cites.
LISP 1.5 Programmer’s Manual
J. McCarthy · 1962
Earlier work this paper cites.
Translator writing systems
J. Feldman and D. Gries · 1968
Earlier work this paper cites.
A systolic array optimizing compiler
M. Lam · 1989
Earlier work this paper cites.
A bridging model for parallel computation
L. Valiant · 1990
Earlier work this paper cites.
The Python programming language, 1994
G. van Rossum et al · 1994
Earlier work this paper cites.
Software pipelining with register allocation and spilling
J. Wang, A. Krall, M. Ertl, and C. Eisenbeis · 1994
Earlier work this paper cites.
POOMA: A Framework for Scientific Simulation on Parallel Architectures
J. Reynders, P. Hinker, J. Cummings, S. Atlas, S. Banerjee, W. Humphrey, K. Keahey, M. Srikant, and M. Tholburn · 1996
Earlier work this paper cites.
Will C++ be faster than Fortran?
T. L. Veldhuizen and M. E. Jernigan · 1997
Earlier work this paper cites.
Independent component filters of natural images compared with simple cells in primary visual cortex
J. H. van Hateren and A. van der Schaaf · 1998
Earlier work this paper cites.
SciPy: Open source scientific tools for Python, 2001–
E. Jones, T. Oliphant, P. Peterson, et al · 2001
Earlier work this paper cites.
Optimizing compilers for modern architectures: a dependence-based approach
K. Kennedy and J. Allen · 2001
Earlier work this paper cites.
Automated empirical optimizations of software and the ATLAS project
R. C. Whaley, A. Petitet, and J. J. Dongarra · 2001
Earlier work this paper cites.
Nodal High-Order Methods on Unstructured Grids: I. Time-Domain Solution of Maxwell’s Equations
J. S. Hesthaven and T. Warburton · 2002
Earlier work this paper cites.
Merrimac: Supercomputing with streams
W. J. Dally, P. Hanrahan, M. Erez, T. J. Knight, F. Labonté, J. H. Ahn, N. Jayasena, U. J. Kapasi, A. Das, and J. Gummaraju · 2003
Earlier work this paper cites.
C++ templates are turing complete
T. L. Veldhuizen · 2003
Earlier work this paper cites.
The graphics card as a stream computer
S. Venkatasubramanian · 2003
Earlier work this paper cites.
Brook for GPUs: stream computing on graphics hardware
I. Buck, T. Foley, D. Horn, J. Sugerman, K. Fatahalian, M. Houston, and P. Hanrahan · 2004
Earlier work this paper cites.
Domain-Specific Program Generation
C. Lengauer, D. Batory, C. Consel, and M. Odersky, editors · 2004
Cited alongside, same era.
Metaprogramming GPUs with Sh
M. McCool and S. D. Toit · 2004
Cited alongside, same era.
MPI for Python
L. Dalcín, R. Paz, and M. Storti · 2005
Cited alongside, same era.
JavaScript at ten years
B. Eich · 2005
Cited alongside, same era.
The design and implementation of FFTW3
M. Frigo and S. G. Johnson · 2005
Cited alongside, same era.
Sage: System for algebra and geometry experimentation
W. Stein and D. Joyner · 2005
Cited alongside, same era.
Programming in Lua
R. Ierusalimschy · 2006
Cited alongside, same era.
The OpenCL 1.0 Specification
K. O. W. Group · 2008
Later among the works it cites.
BSGP: bulk-synchronous GPU programming
Q. Hou, K. Zhou, and B. Guo · 2008
Later among the works it cites.
Nvidia Tesla: A Unified Graphics and Computing Architecture
E. Lindholm, J. Nickolls, S. Oberman, and J. Montrym · 2008
Later among the works it cites.
Larrabee: a many-core x86 architecture for visual computing
L. Seiler, D. Carmean, E. Sprangle, T. Forsyth, M. Abrash, P. Dubey, S. Junkins, A. Lake, J. Sugerman, R. Cavin, R. Espasa, E. Grochowski, T. Juan, and P. Hanrahan · 2008
Later among the works it cites.
A machine vision extension for the Ruby programming language
J. Wedekind, B. Amavasai, K. Dutton, and M. Boissenin · 2008
Later among the works it cites.
Implementing sparse matrix-vector multiplication on throughput-oriented processors
N. Bell and M. Garland · 2009
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implementing an embedded GPU language by combining translation and generation
C. Lejdfors and L. Ohlsson · 2006
Cited alongside, same era.
Data-parallel programming on the Cell BE and the GPU using the RapidMind development platform
M. McCool and R. Inc · 2006
Cited alongside, same era.
Guide to NumPy
T. Oliphant · 2006
Cited alongside, same era.
A domain specific embedded language in C++ for automatic differentiation, projection, integration and variational formulations
C. Prud’homme · 2006
Cited alongside, same era.
Accelerator: using data parallelism to program GPUs for general-purpose uses
D. Tarditi, S. Puri, and J. Oglesby · 2006
Cited alongside, same era.
/hi/cuda: a high-level directive-based language for gpu programming
T. D. Han and T. S. Abdelrahman · 2009
Closest in time.
The CodePy C Code Generation Library, 2009
A. Klöckner · 2009
Closest in time.
Nodal discontinuous Galerkin methods on graphics processors
A. Klöckner, T. Warburton, J. Bridge, and J. Hesthaven · 2009
Closest in time.
Python Scripting for Computational Science
H. P. Langtangen · 2009
Closest in time.
NVIDIA CUDA 2.2 Compute Unified Device Architecture Programming Guide
Nvidia Corporation · 2009
Closest in time.
A high-throughput screening approach to discovering good forms of biologically inspired visual representation
N. Pinto, D. Doukhan, J. DiCarlo, and D. Cox · 2009
Closest in time.
The Jinja 2 Templating Engine, 2009
A. Ronacher · 2009
Closest in time.
JCUDA: a programmer-friendly interface for accelerating Java programs with CUDA
Y. Yan, M. Grossman, and V. Sarkar · 2009
Closest in time.
GPU-accelerated synthetic aperture radar backprojection in CUDA
A. R. Fasih and T. D. R. Hartley · 2010
Closest in time.
Evaluating the Invariance Properties of a Successful Biologically-Inspired Face Recognition System
N. Pinto and D. Cox · 2010
Closest in time.
Copperhead: Compiling an embedded data parallel language
B. Catanzaro, M. Garland, and K. Keutzer · 2011
Closest in time.
"cl.oquence: High-level language abstractions for low-level heterogeneous computing", 2011
C. Omar · 2011
Closest in time.