Fetching the paper…
Reading the bibliography…
An accurate prediction of scheduling and execution of instruction streams is a necessary prerequisite for predicting the in-core performance behavior of throughput-bound loop kernels on out-of-order processor architectures.
H. T. Kung, “Memory Requirements for Balanced Computer Architectures,” in Proceedings of the 13th Annual International Symposium on Computer Architecture , ser. ISCA ’86. Los Alamitos, CA, USA: IEEE Computer Society Press, 1986, pp. 49–54, doi: 10.1145/17356.17362
1986
Earlier work this paper cites.
W. Schönauer, Scientific Supercomputing: Architecture and Use of Shared and Distributed Memory Parallel Computers . Self-edition, 2000. [Online]. Available: http://www.rz.uni-karlsruhe.de/~rx03/book
2000
Earlier work this paper cites.
J. Diamond, M. Burtscher, J. D. McCalpin, B. Kim, S. W. Keckler, and J. C. Browne, “Evaluation and optimization of multicore performance bottlenecks in supercomputing applications,” in (IEEE ISPASS) IEEE International Symposium on Performance Analysis of Systems and Software , April 2011, pp. 32–43
2011
Earlier work this paper cites.
N. Binkert, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. Vaish, M. D. Hill, D. A. Wood, B. Beckmann, G. Black, and et al., “The gem5 simulator,” ACM SIGARCH Computer Architecture News , vol. 39, no. 2, p. 1, 8 2011, doi: 10.1145/2024716.2024718. [Online]. Available: http://dx.doi.org/10.1145/2024716.2024718
2011
Earlier work this paper cites.
A. Patel, F. Afram, and K. Ghose, “Marss-x86: A qemu-based micro-architectural and systems simulator for x86 multicore processors,” in 1st International Qemu Users’ Forum , 2011, pp. 29–30
2011
Earlier work this paper cites.
(2017, 8) Software Optimization Guide for AMD Family 17h Processors. [Online]. Available: https://developer.amd.com/wordpress/media/2013/12/55723_SOG_Fam_17h_Processors_3.00.pdf
2013
Earlier work this paper cites.
D. Sanchez and C. Kozyrakis, “ZSim: Fast and Accurate Microarchitectural Simulation of Thousand-Core Systems,” Proceedings of the 40th Annual International Symposium on Computer Architecture - ISCA ’13 , 2013. [Online]. Available: http://dx.doi.org/10.1145/2485922.2485963
2013
Cited alongside, same era.
A. S. Charif-Rubial, E. Oseret, J. Noudohouenou, W. Jalby, and G. Lartigue, “CQA: A code quality analyzer tool at binary level,” in 2014 21st International Conference on High Performance Computing (HiPC) , Dec 2014, pp. 1–10
2014
Cited alongside, same era.
H. Stengel, J. Treibig, G. Hager, and G. Wellein, “Quantifying Performance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory Model,” in Proceedings of the 29th ACM International Conference on Supercomputing , ser. ICS ’15. New York, NY, USA: ACM, 2015, pp. 207–216, doi: 10.1145/2751205.2751240
2015
Cited alongside, same era.
J. Laukemann. (2017, 12) OSACA – Open Source Architecture Code Analyzer. [Online]. Available: https://github.com/RRZE-HPC/OSACA
2017
Later among the works it cites.
(2018, 4) Instruction tables. [Online]. Available: http://www.agner.org/optimize/instruction_tables.pdf
2018
Closest in time.
2018
Closest in time.
J. Laukemann, “Design and Implemention of a Framework for Predicting Instruction Throughput,” Bachelor’s Thesis, 2017. [Online]. Available: https://hpc.fau.de/files/2018/08/Laukemann_Jan_Design_and_Implementation_For_a_Framework_Predicting_Instruction_Throughput.pdf
2018
Closest in time.
J. Hofmann. (2018, 1) ibench – Measure Instruction Latency and Throughput. [Online]. Available: https://github.com/hofm/ibench
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Lo, S. Williams, B. Van Straalen, T. J. Ligocki, M. J. Cordery, N. J. Wright, M. W. Hall, and L. Oliker, “Roofline Model Toolkit: A Practical Tool for Architectural and Program Analysis,” in High Performance Computing Systems. Performance Modeling, Benchmarking, and Simulation , ser. Lecture Notes in Computer Science, S. A. Jarvis, S. A. Wright, and S. D. Hammond, Eds., vol. 8966. Springer International Publishing, 2015, pp. 129–148, doi: 10.1007/978-3-319-17248-4_7
2015
Cited alongside, same era.
J. Hammer, J. Eitzinger, G. Hager, and G. Wellein, “Kerncraft: A Tool for Analytic Performance Modeling of Loop Kernels,” in Tools for High Performance Computing 2016 , C. Niethammer, J. Gracia, T. Hilbrich, A. Knüpfer, M. M. Resch, and W. E. Nagel, Eds. Cham: Springer International Publishing, 2017, pp. 1–22, doi: 10.1007/978-3-319-56702-0_1
2017
Cited alongside, same era.
(2017, 11) Intel Architecture Code Analyzer. [Online]. Available: https://software.intel.com/en-us/articles/intel-architecture-code-analyzer
2017
Cited alongside, same era.
“Artifact description: Automated instruction stream throughput prediction for intel and amd microarchitectures.” [Online]. Available: https://github.com/RRZE-HPC/pmbs2018-paper-artifact/
Cited in the paper.
Intel 64 and IA-32 Architectures Optimization Reference Manual. [Online]. Available: https://software.intel.com/en-us/download/intel-64-and-ia-32-architectures-optimization-reference-manual
Cited in the paper.
M. Clark. A New X86 Core Architecture for the Next Generation of Computing. [Online]. Available: http://www.hotchips.org/wp-content/uploads/hc_archives/hc28/HC28.23-Tuesday-Epub/HC28.23.90-High-Perform-Epub/HC28.23.930-X86-core-MikeClark-AMD-final_v2-28.pdf
Cited in the paper.
D. Andric. [RFC] llvm-mca: a static performance analysis tool. [Online]. Available: http://llvm.1065342.n5.nabble.com/llvm-dev-RFC-llvm-mca-a-static-performance-analysis-tool-td117477.html
Cited in the paper.
llvm-exegesis – LLVM Machine Instruction Benchmark. [Online]. Available: https://llvm.org/docs/CommandGuide/llvm-exegesis.html
Cited in the paper.
J. Hammer, G. Hager, and G. Wellein, “OoO Instruction Benchmarking Framework on the Back of Dragons,” SC18 SRC Poster (in review)
Cited in the paper.
2018
Closest in time.