Fetching the paper…
Reading the bibliography…
GPU accelerators have become an important backbone for scientific high performance computing, and the performance advances obtained from adopting new GPU hardware are significant.
Templates for the Solution of Linear Systems: Building Blocks for Iterative Methods, 2nd Edition
Richard. Barrett, Michael. Berry, Tony. F. Chan, James. Demmel, June. Donato, Jack. Dongarra, Viktor. Eijkhout, Roldan. Pozo, Charles. Romine, and Henk. Van der Vorst · 1994
Earlier work this paper cites.
Estimating an eigenvector by the power method with a random start
Gianna M. Del Corso · 1997
Earlier work this paper cites.
The PageRank citation ranking: Bringing order to the Web
Lawrence. Page, Sergey. Brin, Rajeev. Motwani, and Terry. Winograd · 1998
Earlier work this paper cites.
Implementing sparse matrix-vector multiplication on throughput-oriented processors
Nathan Bell and Michael Garland · 2009
Earlier work this paper cites.
Roofline: An Insightful Visual Performance Model for Multicore Architectures
Samuel Williams, Andrew Waterman, and David Patterson · 2009
Earlier work this paper cites.
Google’s PageRank and Beyond: The Science of Search Engine Rankings
Amy N. Langville and Carl D. Meyer · 2012
Earlier work this paper cites.
Implementing a Sparse Matrix Vector Product for the SELL-C/SELL-C- σ \sigma formats on NVIDIA GPUs
Hartwig. Anzt, Stanimire. Tomov, and Jack. Dongarra · 2014
Earlier work this paper cites.
A unified sparse matrix data format for efficient general sparse matrix-vector multiplication on modern processors with wide SIMD units
Moritz Kreutzer, Georg Hager, Gerhard Wellein, Holger Fehske, and Alan R. Bishop · 2014
Earlier work this paper cites.
Optimizing sparse matrix operations on gpus using merge path
Steven. Dalton, Sean. Baxter, Duane. Merrill, Luke. Olson, and Michael. Garland · 2015
Cited alongside, same era.
High-performance and scalable GPU graph traversal
Duane Merrill, Michael Garland, and Andrew S. Grimshaw · 2015
Cited alongside, same era.
Gpu-stream v2.0: Benchmarking the achievable memory bandwidth of many-core processors across diverse parallel programming models
Tom Deakin, James Price, Matt Martineau, and Simon McIntosh-Smith · 2016
Cited alongside, same era.
A note on performance profiles for benchmarking software
Nicholas Gould and Jennifer Scott · 2016
Cited alongside, same era.
Merge-based parallel sparse matrix-vector multiplication
Duane Merrill and Michael Garland · 2016
Cited alongside, same era.
Preconditioned Krylov solvers on GPUs
Hartwig Anzt, Mark Gates, Jack Dongarra, Moritz Kreutzer, Gerhard Wellein, and Martin Köhler · 2017
Balanced csr sparse matrix-vector product on graphics processors
Goran Flegar and Enrique S. Quintana-Ortí · 2017
Later among the works it cites.
Matrix Collection
SuiteSparse · 2018
Later among the works it cites.
Adaptive sparse tiling for sparse matrix multiplication
Changwan Hong, Aravind Sukumaran-Rajam, Israt Nisa, Kunal Singh, and P. Sadayappan · 2019
Later among the works it cites.
Ginkgo: A modern linear operator algebra framework for high performance computing, 2020
Hartwig Anzt, Terry Cojean, Goran Flegar, Fritz Göbel, Thomas Grützmacher, Pratik Nayak, Tobias Ribizel, Yuhsiang Mike Tsai, and Enrique S. Quintana-Ortí · 2020
Closest in time.
Load-Balancing Sparse Matrix Vector Product Kernels on GPUs
Hartwig Anzt, Terry Cojean, Chen Yen-Chen, Jack Dongarra, Goran Flegar, Pratik Nayak, Stanimire Tomov, Yuhsiang M. Tsai, and Weichung Wang · 2020
Closest in time.
CUDA 11.0 Release Notes
NIVIDA · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Overcoming load imbalance for irregular sparse matrices
Goran Flegar and Hartwig Anzt · 2017
Cited alongside, same era.
The Top 500 List, http://www.top.org/
Cited in the paper.
Closest in time.
NVIDIA A100 Tensor Core GPU Architecture
Nvidia · 2020
Closest in time.