Fetching the paper…
Reading the bibliography…
GPUs are readily available in cloud computing and personal devices, but their use for data processing acceleration has been slowed down by their limited integration with common programming languages such as Python or Java.
J. M. Kleinberg, “Authoritative sources in a hyperlinked environment,” Journal of the ACM (JACM) , vol. 46, no. 5, pp. 604–632, 1999
1999
Earlier work this paper cites.
S. Che, M. Boyer, J. Meng, D. Tarjan, J. W. Sheaffer, S.-H. Lee, and K. Skadron, “Rodinia: A benchmark suite for heterogeneous computing,” in 2009 IEEE international symposium on workload characterization (IISWC) . Ieee, 2009, pp. 44–54
2009
Earlier work this paper cites.
J. E. Stone, D. Gohara, and G. Shi, “Opencl: A parallel programming standard for heterogeneous computing systems,” Computing in science & engineering , vol. 12, no. 3, p. 66, 2010
2010
Earlier work this paper cites.
V. T. Ravi, M. Becchi, G. Agrawal, and S. Chakradhar, “Supporting gpu sharing in cloud environments with a transparent runtime consolidation framework,” in Proceedings of the 20th international symposium on High performance distributed computing , 2011, pp. 217–228
2011
Earlier work this paper cites.
C. Wimmer and T. Würthinger, “Truffle: a self-optimizing runtime system,” in Proceedings of the 3rd annual conference on Systems, programming, and applications: software for humanity . ACM, 2012
2012
Earlier work this paper cites.
A. Klöckner, N. Pinto, Y. Lee, B. Catanzaro, P. Ivanov, and A. Fasih, “PyCUDA and PyOpenCL: A Scripting-Based Approach to GPU Run-Time Code Generation,” Parallel Computing , vol. 38, 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
T. Würthinger, C. Wimmer, A. Wöß, L. Stadler, G. Duboscq, C. Humer, G. Richards, D. Simon, and M. Wolczko, “One vm to rule them all,” in Proceedings of the 2013 ACM international symposium on New ideas, new paradigms, and reflections on programming & software , 2013
2013
Earlier work this paper cites.
G. Duboscq, L. Stadler, T. Würthinger, D. Simon, C. Wimmer, and H. Mössenböck, “Graal ir: An extensible declarative intermediate representation,” in Proceedings of the Asia-Pacific Programming Languages and Compilers Workshop , 2013
2013
Earlier work this paper cites.
T. Gautier, J. V. Lima, N. Maillard, and B. Raffin, “Xkaapi: A runtime system for data-flow task programming on heterogeneous architectures,” in 2013 IEEE 27th International Symposium on Parallel and Distributed Processing . IEEE, 2013, pp. 1299–1308
2013
Earlier work this paper cites.
J. Luitjens, “Faster parallel reductions on kepler,” developer.nvidia.com/blog/faster-parallel-reductions-kepler
2014
Cited alongside, same era.
J. Luitjens, “Cuda streams: Best practices and common pitfalls,” in GPU Techonology Conference , 2015
2015
Cited alongside, same era.
Q. Chen, H. Yang, J. Mars, and L. Tang, “Baymax: Qos awareness and increased utilization for non-preemptive accelerators in warehouse scale computers,” ACM SIGPLAN Notices , vol. 51, no. 4, pp. 681–696, 2016
2016
Cited alongside, same era.
R. Mayer, C. Mayer, and L. Laich, “The tensorflow partitioning and scheduling problem: it’s the critical path!” in Proceedings of the 1st Workshop on Distributed Infrastructures for Deep Learning , 2017
2017
Cited alongside, same era.
C.-H. Hong, I. Spence, and D. S. Nikolopoulos, “Gpu virtualization and scheduling methods: A comprehensive survey,” ACM Computing Surveys (CSUR) , vol. 50, no. 3, pp. 1–37, 2017
J. Fumero, M. Papadimitriou, F. S. Zakkak, M. Xekalaki, J. Clarkson, and C. Kotselidis, “Dynamic application reconfiguration on heterogeneous hardware,” in Proceedings of the 15th ACM SIGPLAN/SIGOPS International Conference on Virtual Execution Environments , 2019
2019
Later among the works it cites.
F. Guo, Y. Li, J. C. Lui, and Y. Xu, “Dcuda: Dynamic gpu scheduling with live migration support,” in Proceedings of the ACM Symposium on Cloud Computing , 2019, pp. 114–125
2019
Later among the works it cites.
“Cuda graphs,” https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__GRAPH.html
2020
Closest in time.
Y. Xu, L. Liu, and Z. Ding, “Dag-aware joint task scheduling and cache management in spark clusters,” in 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2020, pp. 378–387
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Y. Wen and M. F. O’Boyle, “Merge or separate? multi-job scheduling for opencl kernels on cpu/gpu platforms,” in Proceedings of the general purpose GPUs , 2017, pp. 22–31
2017
Cited alongside, same era.
L. Marchal, H. Nagy, B. Simon, and F. Vivien, “Parallel scheduling of dags under memory constraints,” in 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2018
2018
Cited alongside, same era.
J. Clarkson, J. Fumero, M. Papadimitriou, F. S. Zakkak, M. Xekalaki, C. Kotselidis, and M. Luján, “Exploiting high-performance heterogeneous hardware for java programs using graal,” in Proceedings of the 15th International Conference on Managed Languages & Runtimes , 2018, pp. 1–13
2018
Cited alongside, same era.
M. Y. Özkaya, A. Benoit, B. Uçar, J. Herrmann, and Ü. V. Çatalyürek, “A scalable clustering-based task scheduler for homogeneous processors using dag partitioning,” in 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2019, pp. 155–165
2019
Cited alongside, same era.
R. Mueller and L. Stadler, “Grcuda,” github.com/NVIDIA/grcuda
Cited in the paper.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: A system for large-scale machine learning,” in 12th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 16) , 2016, pp. 265–283
Cited in the paper.
“Cuda gaussian blur,” github.com/harrytang/cuda-gaussian-blur
Cited in the paper.
A. Marchetti-Spaccamela, N. Megow, J. Schlöter, M. Skutella, and L. Stougie, “On the complexity of conditional dag scheduling in multiprocessor systems,” in 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2020, pp. 1061–1070
2020
Closest in time.
A. Gray, “Getting started with cuda graphs,” developer.nvidia.com/blog/cuda-graphs
2020
Closest in time.
B. Qiao, O. Reiche, J. Teich, and F. Hannig, “Unveiling kernel concurrency in multiresolution filters on gpus with an image processing dsl,” in Proceedings of the 13th Annual Workshop on General Purpose Processing using Graphics Processing Unit , 2020, pp. 11–20
2020
Closest in time.
J. Jung, D. Park, Y. Do, J. Park, and J. Lee, “Overlapping host-to-device copy and computation using hidden unified memory,” in Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2020, pp. 321–335
2020
Closest in time.
“Black & scholes option pricing,” docs.nvidia.com/cuda/cuda-samples/index.html#black-scholes-option-pricing
2020
Closest in time.