Fetching the paper…
Reading the bibliography…
Porting code from CPU to GPU is costly and time-consuming; Unless much time is invested in development and optimization, it is not obvious, a priori, how much speed-up is achievable or how much room is left for improvement.
D. H. Bailey, E. Barszcz, J. T. Barton, D. S. Browning, R. L. Carter, L. Dagum, R. A. Fatoohi, P. O. Frederickson, T. A. Lasinski, R. S. Schreiber, et al
1991
Earlier work this paper cites.
R. H. Saavedra and A. J. Smith, “Analysis of benchmark characteristics and benchmark performance prediction,” TOCS
1996
Earlier work this paper cites.
J. J. Yi, D. J. Lilja, and D. M. Hawkins, “A statistically rigorous approach for improving simulation methodology,” in HPCA
2003
Earlier work this paper cites.
H. Vandierendonck and K. De Bosschere, “Many benchmarks stress the same bottlenecks,” in Workshop on Computer Architecture Evaluation Using Commercial Workloads
2004
Earlier work this paper cites.
K. Hoste, A. Phansalkar, L. Eeckhout, A. Georges, L. K. John, and K. D. Bosschere, “Performance prediction based on inherent program similarity,” in PACT
2006
Earlier work this paper cites.
B. C. Lee and D. M. Brooks, “Accurate and efficient regression modeling for microarchitectural performance and power prediction,” in ASPLOS
2006
Earlier work this paper cites.
E. Schweitz, R. Lethin, A. Leung, and B. Meister, “R-stream: A parametric high level compiler,” HPEC
2006
Earlier work this paper cites.
B. Lee and D. Brooks, “Illustrative design space studies with microarchitectural regression models,” in HPCA
2007
Cited alongside, same era.
C. Dubach, J. Cavazos, B. Franke, G. Fursin, M. F. O’Boyle, and O. Temam, “Fast compiler optimisation evaluation using code-feature based performance prediction,” in CF
2007
Cited alongside, same era.
S.-Z. Ueng, M. Lathara, S. S. Baghsorkhi, and W. H. Wen-mei, “CUDA-lite: Reducing gpu programming complexity,” in LCPC
2008
Cited alongside, same era.
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Commun. ACM
2009
Cited alongside, same era.
M. Kulkarni, M. Burtscher, C. Casçaval, and K. Pingali, “Lonestar: A suite of parallel irregular programs,” in ISPASS
2009
Cited alongside, same era.
W. Wu and B. C. Lee, “Inferred models for dynamic and sparse hardware-software spaces,” in MICRO
2012
Later among the works it cites.
D. Mikushin and N. Likhogrud, “KERNELGEN–a toolchain for automatic gpu-centric applications porting,” 2012
2012
Later among the works it cites.
M. R. Meswani, L. Carrington, D. Unat, A. Snavely, S. Baden, and S. Poole, “Modeling and predicting performance of high performance computing applications on hardware accelerators,” IJHPC
2013
Later among the works it cites.
PhD thesis, Princeton University, 2013
T. B. Jablin, Automatic Parallelization for GPUs · 2013
Later among the works it cites.
T. Hoshino, N. Maruyama, S. Matsuoka, and R. Takaki, “Cuda vs openacc: Performance case studies with kernel benchmarks and a memory-bound cfd application,” in Cluster, Cloud and Grid Computing (CCGrid), 2013 13th IEEE/ACM International Symposium on
2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Meng, V. Morozov, K. Kumaran, V. Vishwanath, and T. Uram, “Grophecy: Gpu performance projection from cpu code skeletons,” in SC
2011
Cited alongside, same era.
C. Nugteren and H. Corporaal, “The boat hull model: adapting the roofline model to enable performance prediction for parallel computing,” in PPOPP ’12
2012
Cited alongside, same era.
I. Baldini, S. J. Fink, and E. Altman, “Predicting gpu performance from cpu runs using machine learning,” in SBAC-PAD ’14
Cited in the paper.
S. Che, M. Boyer, J. Meng, D. Tarjan, J. W. Sheaffer, S.-H. Lee, and K. Skadron, “Rodinia: A benchmark suite for heterogeneous computing,” in IISWC ’09
Cited in the paper.
http://hpcgpu.codeplex.com
L. L. Pilla, “NAS Parallel Benchmarks CUDA version.”
Cited in the paper.
http://docs.nvidia.com/cuda/cuda-c-best-practices-guide/
“Cuda Toolkit Documentation.”
Cited in the paper.
A. P. L. K. John, “Performance prediction using program similarity,”
Cited in the paper.
Later among the works it cites.
W. Jia, K. A. Shaw, and M. Martonosi, “Starchart: hardware and software optimization using recursive partitioning regression trees,” in PACT
2013
Later among the works it cites.
N. Ardalani, C. Lestourgeon, K. Sankaralingam, and X. Zhu, “Cross-architecture performance prediction (xapp): Using cpu code to predict gpu performance,” in MICRO
2015
Later among the works it cites.