Fetching the paper…
Reading the bibliography…
Useful models of loop kernel runtimes on out-of-order architectures require an analysis of the in-core performance behavior of instructions and their dependencies.
U. Manber, “Introduction to algorithms - a creative approach,” 1989
1989
Earlier work this paper cites.
R. Barrett, M. Berry, T. F. Chan, J. Demmel, J. Donato, J. Dongarra, V. Eijkhout, R. Pozo, C. Romine, and H. V. der Vorst, Templates for the Solution of Linear Systems: Building Blocks for Iterative Methods, 2nd Edition . Philadelphia, PA: SIAM, 1994
1994
Earlier work this paper cites.
S. Williams, A. Waterman, and D. Patterson, “Roofline: An insightful visual performance model for multicore architectures,” Commun. ACM , vol. 52, no. 4, pp. 65–76, 2009
2009
Earlier work this paper cites.
N. Binkert, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. Vaish, M. D. Hill, D. A. Wood, B. Beckmann, G. Black, and et al., “The gem5 simulator,” ACM SIGARCH Computer Architecture News , vol. 39, no. 2, p. 1, 8 2011, doi: 10.1145/2024716.2024718. [Online]. Available: http://dx.doi.org/10.1145/2024716.2024718
2011
Earlier work this paper cites.
A. Patel, F. Afram, and K. Ghose, “Marss-x86: A qemu-based micro-architectural and systems simulator for x86 multicore processors,” in 1st International Qemu Users’ Forum , 2011, pp. 29–30
2011
Earlier work this paper cites.
D. Sanchez and C. Kozyrakis, “ZSim: Fast and Accurate Microarchitectural Simulation of Thousand-Core Systems,” Proceedings of the 40th Annual International Symposium on Computer Architecture - ISCA ’13 , 2013. [Online]. Available: http://dx.doi.org/10.1145/2485922.2485963
2013
Earlier work this paper cites.
A. S. Charif-Rubial, E. Oseret, J. Noudohouenou, W. Jalby, and G. Lartigue, “CQA: A code quality analyzer tool at binary level,” in 2014 21st International Conference on High Performance Computing (HiPC) , Dec 2014, pp. 1–10, doi: 10.1109/HiPC.2014.7116904
2014
Earlier work this paper cites.
H. Stengel, J. Treibig, G. Hager, and G. Wellein, “Quantifying Performance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory Model,” in Proceedings of the 29th ACM International Conference on Supercomputing , ser. ICS ’15. New York, NY, USA: ACM, 2015, pp. 207–216, doi: 10.1145/2751205.2751240
2015
Cited alongside, same era.
Y. Lo, S. Williams, B. Van Straalen, T. J. Ligocki, M. J. Cordery, N. J. Wright, M. W. Hall, and L. Oliker, “Roofline Model Toolkit: A Practical Tool for Architectural and Program Analysis,” in High Performance Computing Systems. Performance Modeling, Benchmarking, and Simulation , ser. Lecture Notes in Computer Science, S. A. Jarvis, S. A. Wright, and S. D. Hammond, Eds., vol. 8966. Springer International Publishing, 2015, pp. 129–148, doi: 10.1007/978-3-319-17248-4_7
2015
Cited alongside, same era.
V. Palomares, D. C. Wong, D. J. Kuck, and W. Jalby, “Evaluating out-of-order engine limitations using uop flow simulation,” in Tools for High Performance Computing 2015 , A. Knüpfer, T. Hilbrich, C. Niethammer, J. Gracia, W. E. Nagel, and M. M. Resch, Eds. Cham: Springer International Publishing, 2016, pp. 161–181, doi: 10.1007/978-3-319-39589-0_13
(2018, 4) Instruction tables. [Online]. Available: http://www.agner.org/optimize/instruction_tables.pdf
2018
Later among the works it cites.
J. Hammer, G. Hager, and G. Wellein, “OoO Instruction Benchmarking Framework on the Back of Dragons,” 2018, SC18 ACM SRC Poster. [Online]. Available: https://sc18.supercomputing.org/proceedings/src_poster/src_poster_pages/spost115.html
2018
Later among the works it cites.
J. Hofmann. (2018, 1) ibench – Measure Instruction Latency and Throughput. [Online]. Available: https://github.com/hofm/ibench
2018
Later among the works it cites.
“Artifact description: Automatic throughput and critical path analysis of x86 and arm assembly kernels.” [Online]. Available: https://github.com/RRZE-HPC/OSACA-CP-2019
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
J. Hammer, J. Eitzinger, G. Hager, and G. Wellein, “Kerncraft: A Tool for Analytic Performance Modeling of Loop Kernels,” in Tools for High Performance Computing 2016 , C. Niethammer, J. Gracia, T. Hilbrich, A. Knüpfer, M. M. Resch, and W. E. Nagel, Eds. Cham: Springer International Publishing, 2017, pp. 1–22, doi: 10.1007/978-3-319-56702-0_1
2017
Cited alongside, same era.
(2017, 11) Intel Architecture Code Analyzer. [Online]. Available: https://software.intel.com/en-us/articles/intel-architecture-code-analyzer
2017
Cited alongside, same era.
J. Laukemann. (2017, 12) OSACA – Open Source Architecture Code Analyzer. [Online]. Available: https://github.com/RRZE-HPC/OSACA
2017
Cited alongside, same era.
J. Laukemann, J. Hammer, J. Hofmann, G. Hager, and G. Wellein, “Automated instruction stream throughput prediction for intel and amd microarchitectures,” in 2018 IEEE/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS) , Nov 2018, pp. 121–131
2018
Cited alongside, same era.
D. Andric. [RFC] llvm-mca: a static performance analysis tool. [Online]. Available: http://llvm.1065342.n5.nabble.com/llvm-dev-RFC-llvm-mca-a-static-performance-analysis-tool-td117477.html
Cited in the paper.
llvm-exegesis – LLVM Machine Instruction Benchmark. [Online]. Available: https://llvm.org/docs/CommandGuide/llvm-exegesis.html
Cited in the paper.
C. Mendis, A. Renda, S. Amarasinghe, and M. Carbin, “Ithemal: Accurate, portable and fast basic block throughput estimation using deep neural networks,” in Proceedings of the 36th International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. Long Beach, California, USA: PMLR, Jun 2019, pp. 4505–4515. [Online]. Available: http://proceedings.mlr.press/v97/mendis19a.html
2019
Closest in time.
A. Abel and J. Reineke, “uops.info: Characterizing latency, throughput, and port usage of instructions on intel microarchitectures,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’19. New York, NY, USA: ACM, 2019, pp. 673–686. [Online]. Available: http://doi.acm.org/10.1145/3297858.3304062
2019
Closest in time.