Fetching the paper…
Reading the bibliography…
GPUs and other accelerators are popular devices for accelerating compute-intensive, parallelizable applications.
N. Solntseff and A. Yezerski, “A survey of extensible programming languages,” Annual review in automatic programming , vol. 7, 1974
1974
Earlier work this paper cites.
J. J. Dongarra, J. Du Croz, S. Hammarling, and I. S. Duff, “A set of level 3 basic linear algebra subprograms,” ACM Trans. Mathematical Software , vol. 16, no. 1, pp. 1–17, 1990
1990
Earlier work this paper cites.
E. Anderson, Z. Bai, C. Bischof, L. S. Blackford, J. Demmel, J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, A. McKenney et al. , LAPACK Users’ guide . SIAM, 1999
1999
Earlier work this paper cites.
D. M. Ciemiewicz, “What do you ‘mean’? Revisiting statistics for web response time measurements,” in Proc. Conf. Computer Measurement Group , 2001, pp. 385–396
2001
Earlier work this paper cites.
C. Lattner and V. Adve, “LLVM: A compilation framework for lifelong program analysis & transformation,” in Proc. Int Symp. Code Generation and Optimization , 2004, pp. 75–86
2004
Earlier work this paper cites.
J. R. Mashey, “War of the benchmark means: time for a truce,” ACM SIGARCH Computer Architecture News , vol. 32, no. 4, 2004
2004
Earlier work this paper cites.
D. Zingaro, “Modern extensible languages,” SQRL Report , vol. 47, p. 16, 2007
2007
Earlier work this paper cites.
C. Lejdfors and L. Ohlsson, “PyGPU: A high-level language for high-speed image processing,” in Int. Conf. Applied Computing 2007 . IADIS, 2007, pp. 66–81
2007
Earlier work this paper cites.
NVIDIA, “cuBLAS: Dense linear algebra on GPUs,” 2008. [Online]. Available: https://developer.nvidia.com/cublas
2008
Earlier work this paper cites.
S. Che, M. Boyer, J. Meng, D. Tarjan, J. W. Sheaffer, S.-H. Lee, and K. Skadron, “Rodinia: A benchmark suite for heterogeneous computing,” in Int. Symp. Workload Characterization , 2009, pp. 44–54
2009
Earlier work this paper cites.
S. Sarkar, T. Majumder, A. Kalyanaraman, and P. P. Pande, “Hardware accelerators for biocomputing: A survey,” in Proc. Int. Symp. Circuits and Systems . IEEE, 2010, pp. 3789–3792
2010
Earlier work this paper cites.
J. Hoberock and N. Bell, “Thrust: A parallel template library,” 2010. [Online]. Available: https://developer.nvidia.com/thrust
2010
Earlier work this paper cites.
K. Rupp, F. Rudolf, and J. Weinbub, “ViennaCL: A high level linear algebra library for GPUs and multi-core CPUs,” in Intl. Workshop GPUs and Scientific Applications , 2010, pp. 51–56
2010
Earlier work this paper cites.
G. Pratx and L. Xing, “GPU computing in medical physics: A review,” Medical physics , vol. 38, no. 5, pp. 2685–2697, 2011
2011
Earlier work this paper cites.
B. Catanzaro, M. Garland, and K. Keutzer, “Copperhead: Compiling an embedded data parallel language,” ACM SIGPLAN Notices , vol. 46, no. 8, pp. 47–56, 2011
2011
Earlier work this paper cites.
M. M. Chakravarty, G. Keller, S. Lee, T. L. McDonell, and V. Grover, “Accelerating Haskell array codes with multicore GPUs,” in Proc. 6th Workshop Declarative Aspects of Multicore Programming . ACM, 2011, pp. 3–14
2011
Earlier work this paper cites.
E. Holk et al. , “GPU programming in Rust: Implementing high-level abstractions in a systems-level language,” in Parallel and Distributed Processing Symp. Workshops & PhD Forum , 2013
2011
Cited alongside, same era.
P. Du, R. Weber, P. Luszczek, S. Tomov, G. Peterson, and J. Dongarra, “From CUDA to OpenCL: Towards a performance-portable solution for multi-platform gpu programming,” Parallel Computing , vol. 38, no. 8, pp. 391–407, 2012
2012
Cited alongside, same era.
J. Malcolm, P. Yalamanchili, C. McClanahan, V. Venugopalakrishnan, K. Patel, and J. Melonakos, “Arrayfire: a GPU acceleration platform,” in Proc. SPIE , vol. 8403, 2012, pp. 84 030A–1
2012
Cited alongside, same era.
2012
Cited alongside, same era.
T. Besard, B. De Sutter, A. Frías-Velázquez, and W. Philips, “Case study of multiple trace transform implementations,” Int. J. High Performance Computing Applications , vol. 29, no. 4, pp. 489–505, 2015
2015
Later among the works it cites.
J. Bezanson, “Why is Julia fast? Can it be faster?” 2015, JuliaCon India
2015
Later among the works it cites.
J. Luitjens. (2015) Faster parallel reductions on Kepler. [Online]. Available: https://devblogs.nvidia.com/parallelforall/faster-parallel-reductions-kepler/
2015
Later among the works it cites.
2015
Later among the works it cites.
C. Kachris and D. Soudris, “A survey on reconfigurable accelerators for cloud computing,” in Int. Conf. Field Programmable Logic and Applications . IEEE, 2016, pp. 1–10
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Rubinsteyn et al. , “Parakeet: A just-in-time parallel accelerator for Python,” in USENIX Conf. Hot Topics in Parallelism , 2012
2012
Cited alongside, same era.
C. Dubach et al. , “Compiling a high-level language for GPUs,” ACM SIGPLAN Notices , vol. 47, no. 6, pp. 1–12, 2012
2012
Cited alongside, same era.
A. Stromme et al. , “Chestnut: A GPU programming language for non-experts,” in Proc. Int. Workshop on Programming Models and Applications for Multicores and Manycores , 2012
2012
Cited alongside, same era.
P. C. Pratt-Szeliga et al. , “Rootbeer: Seamlessly using GPUs from Java,” in Int. Conf. High Performance Computing and Communication , 2012, pp. 375–380
2012
Cited alongside, same era.
S. Aluru and N. Jammula, “A review of hardware acceleration for computational genomics,” IEEE Design & Test , vol. 31, no. 1, 2014
2014
Cited alongside, same era.
N. Markovskiy. (2014, 6) Drop-in acceleration of GNU Octave. NVIDIA. [Online]. Available: https://devblogs.nvidia.com/parallelforall/drop-in-acceleration-gnu-octave/
2014
Cited alongside, same era.
M. Ragan-Kelley, F. Perez, B. Granger, T. Kluyver, P. Ivanov, J. Frederic, and M. Bussonnier, “The Jupyter/IPython architecture: a unified view of computational research, from interactive exploration to communication and publication.” in AGU Fall Meeting Abstracts , 2014
2014
Cited alongside, same era.
J. Weerasinghe, F. Abel, C. Hagleitner, and A. Herkersdorf, “Enabling FPGAs in hyperscale data centers,” in Proc. Int. Conf. Ubiquitous Intelligence and Computing, Autonomic and Trusted Computing, Scalable Computing and Communications . IEEE, 2015, pp. 1078–1086
2015
Cited alongside, same era.
2016
Later among the works it cites.
X. Li, P.-C. Shih, J. Overbey, C. Seals, and A. Lim, “Comparing programmer productivity in OpenACC and CUDA: an empirical investigation,” Int. J. Computer Science, Engineering and Applications , vol. 6, no. 5, pp. 1–15, 2016
2016
Later among the works it cites.
J. Wu, A. Belevich, E. Bendersky, M. Heffernan, C. Leary, J. Pienaar, B. Roune, R. Springer, X. Weng, and R. Hundt, “gpucc: An open-source GPGPU compiler,” in Proc. Int. Symp. Code Generation and Optimization . ACM, 2016, pp. 105–116
2016
Later among the works it cites.
J. Nash, “Inference convergence algorithm in Julia,” 2016. [Online]. Available: https://juliacomputing.com/blog/2016/04/04/inference-convergence.html
2016
Later among the works it cites.
2016
Later among the works it cites.
R. Membarth, O. Reiche, F. Hannig, J. Teich, M. Körner, and W. Eckert, “HIPA cc : A domain-specific language and compiler for image processing,” IEEE Transactions on Parallel and Distributed Systems , vol. 27, no. 1, pp. 210–224, 2016
2016
Later among the works it cites.
Continuum Analytics, “Anaconda Accelerate: GPU-accelerated numerical libraries for Python,” 2017. [Online]. Available: https://docs.anaconda.com/accelerate/
2017
Closest in time.
Julia developers, “CUBLAS.jl: Julia interface to cuBLAS,” 2017. [Online]. Available: https://github.com/JuliaGPU/CUBLAS.jl/
2017
Closest in time.
J. Bezanson, A. Edelman, S. Karpinski, and V. B. Shah, “Julia: A fresh approach to numerical computing,” SIAM Review , vol. 59, no. 1, pp. 65–98, 2017
2017
Closest in time.
M. Innes, “CuArrays.jl: CUDA-accelerated arrays for Julia,” 2017. [Online]. Available: https://github.com/FluxML/CuArrays.jl
2017
Closest in time.
S. G. Johnson. (2017) More dots: Syntactic loop fusion in Julia. [Online]. Available: https://julialang.org/blog/2017/01/moredots
2017
Closest in time.