Fetching the paper…
Reading the bibliography…
Generalized Sparse Matrix-Matrix Multiplication (SpGEMM) is a ubiquitous task in various engineering and scientific applications.
F. G. Gustavson, “Two fast algorithms for sparse matrices: Multiplication and permuted transposition,” ACM Transactions on Mathematical Software (TOMS)
1978
Earlier work this paper cites.
M. O. Rabin and V. V. Vazirani, “Maximum matchings in general graphs through randomization,” Journal of Algorithms
1989
Earlier work this paper cites.
G. Karypis, A. Gupta, and V. Kumar, “A parallel formulation of interior point algorithms,” in Supercomputing’94: Proceedings of the 1994 ACM/IEEE Conference on Supercomputing
1994
Earlier work this paper cites.
S. Itoh, P. Ordejón, and R. M. Martin, “Order-n tight-binding molecular dynamics on parallel computers,” Computer physics communications
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation
1997
Earlier work this paper cites.
PhD thesis, 2000
S. M. Van Dongen, Graph clustering by flow simulation · 2000
Earlier work this paper cites.
L. Zhuo and V. K. Prasanna, “Sparse matrix-vector multiplication on fpgas,” in Proceedings of the 2005 ACM/SIGDA 13th International Symposium on Field-programmable Gate Arrays
2005
Earlier work this paper cites.
J. R. Gilbert, S. Reinhardt, and V. B. Shah, “High-performance graph algorithms from parallel sparse matrices,” in International Workshop on Applied Parallel Computing
2006
Earlier work this paper cites.
G. Penn, “Efficient transitive closure of sparse matrices over closed semirings,” Theoretical Computer Science
2006
Earlier work this paper cites.
J. R. Gilbert, S. Reinhardt, and V. B. Shah, “A unified framework for numerical and combinatorial computing,” Computing in Science & Engineering
2008
Earlier work this paper cites.
G. Goumas, K. Kourtis, N. Anastopoulos, V. Karakasis, and N. Koziris, “Understanding the performance of sparse matrix-vector multiplication,” in 16th Euromicro Conference on Parallel, Distributed and Network-Based Processing (PDP 2008)
2008
Earlier work this paper cites.
Y. Elkurdi, D. Fernández, E. Souleimanov, D. Giannacopoulos, and W. J. Gross, “Fpga architecture and implementation of sparse matrix–vector multiplication for the finite element method,” Computer Physics Communications
2008
Earlier work this paper cites.
S. Asano, T. Maruyama, and Y. Yamaguchi, “Performance comparison of fpga, gpu and cpu in image processing,” in Field programmable logic and applications, 2009. fpl 2009. international conference on
2009
Earlier work this paper cites.
T. M. Chan, “More algorithms for all-pairs shortest paths in weighted graphs,” SIAM Journal on Computing
2010
Earlier work this paper cites.
I. Yamazaki and X. S. Li, “On techniques to improve robustness and scalability of a parallel hybrid linear solver,” in International Conference on High Performance Computing for Computational Science
2010
Earlier work this paper cites.
R. C. Murphy, K. B. Wheeler, B. W. Barrett, and J. A. Ang, “Introducing the graph 500,” Cray Users Group (CUG)
2010
Earlier work this paper cites.
S. Galal and M. Horowitz, “Energy-efficient floating-point unit design,” IEEE Transactions on computers
2010
Earlier work this paper cites.
T. A. Davis and Y. Hu, “The university of florida sparse matrix collection,” ACM Transactions on Mathematical Software (TOMS)
2011
Earlier work this paper cites.
F. Zhang, D. Wu, N. Ao, G. Wang, X. Liu, and J. Liu, “Fast lists intersection with bloom filter using graphics processing units,” in Proceedings of the 2011 ACM Symposium on Applied Computing
2011
Earlier work this paper cites.
N. Ao, F. Zhang, D. Wu, D. S. Stones, G. Wang, X. Liu, J. Liu, and S. Lin, “Efficient parallel lists intersection and index compression algorithms using graphics processing units,” Proceedings of the VLDB Endowment
2011
Cited alongside, same era.
N. Bell, S. Dalton, and L. N. Olson, “Exposing fine-grained parallelism in algebraic multigrid methods,” SIAM Journal on Scientific Computing
2012
Cited alongside, same era.
K. Matam, S. R. K. B. Indarapu, and K. Kothapalli, “Sparse matrix-matrix multiplication on modern architectures,” in 2012 19th International Conference on High Performance Computing
2012
Cited alongside, same era.
N. Bell, S. Dalton, and L. N. Olson, “Exposing fine-grained parallelism in algebraic multigrid methods,” SIAM Journal on Scientific Computing
2012
Cited alongside, same era.
P. Grigoraş, P. Burovskiy, W. Luk, and S. Sherwin, “Optimising sparse matrix vector multiplication for large scale fem problems on fpga,” in Field Programmable Logic and Applications (FPL), 2016 26th International Conference on
2016
Later among the works it cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: efficient inference engine on compressed deep neural network,” in ISCA
2016
Later among the works it cites.
M. Fort, J. A. Sellarès, and N. Valladares, “Intersecting two families of sets on the gpu,” Journal of Parallel and Distributed Computing
2017
Later among the works it cites.
S. Han, J. Kang, H. Mao, Y. Hu, X. Li, Y. Li, D. Xie, H. Luo, S. Yao, Y. Wang, et al
2017
Later among the works it cites.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” in ISCA
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
D. Zou, Y. Dou, S. Guo, and S. Ni, “High performance sparse matrix-vector multiplication on fpga,” IEICE Electronics Express
2013
Cited alongside, same era.
S. Badia, A. F. Martín, and J. Principe, “A highly scalable parallel implementation of balancing domain decomposition by constraints,” SIAM Journal on Scientific Computing
2014
Cited alongside, same era.
J. Leskovec and A. Krevl, “SNAP Datasets: Stanford large network dataset collection.” http://snap.stanford.edu/data
2014
Cited alongside, same era.
Version 0.5.0
S. Dalton, N. Bell, L. Olson, and M. Garland, “Cusp: Generic parallel algorithms for sparse matrix and graph computations,” 2014 · 2014
Cited alongside, same era.
W. Liu and B. Vinter, “An efficient gpu general sparse matrix-matrix multiplication for irregular data,” in 2014 IEEE 28th International Parallel and Distributed Processing Symposium
2014
Cited alongside, same era.
2015
Cited alongside, same era.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in neural information processing systems
2015
Cited alongside, same era.
2017
Later among the works it cites.
S. Pal, J. Beaumont, D.-H. Park, A. Amarnath, S. Feng, C. Chakrabarti, H.-S. Kim, D. Blaauw, T. Mudge, and R. Dreslinski, “Outerspace: An outer product based sparse matrix multiplication accelerator,” in 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA)
2018
Later among the works it cites.
Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “Amc: Automl for model compression and acceleration on mobile devices,” in ECCV
2018
Later among the works it cites.
H. Wang, J. Yang, H.-S. Lee, and S. Han, “Learning to design circuits,” NeurIPS 2018 Machine Learning for Systems Workshop
2018
Later among the works it cites.
J. Cong, Z. Fang, M. Lo, H. Wang, J. Xu, and S. Zhang, “Understanding performance differences of fpgas and gpus,” in 2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
2018
Later among the works it cites.
J. L. Hennessy and D. A. Patterson, “A new golden age for computer architecture: Domain-specific hardware/software co-design, enhanced security, open instruction sets, and agile chip development,” Turing Lecture
2018
Later among the works it cites.
S. Han, L. Zou, and J. X. Yu, “Speeding up set intersections in graph algorithms using simd instructions,” in Proceedings of the 2018 International Conference on Management of Data
2018
Later among the works it cites.
H. Mao, P. Negi, A. Narayan, H. Wang, J. Yang, H. Wang, R. Marcus, r. addanki, M. Khani Shirkoohi, S. He, V. Nathan, F. Cangialosi, S. Venkatakrishnan, W.-H. Weng, S. Han, T. Kraska, and D. Alizadeh, “Park: An open platform for learning-augmented computer systems,” in Advances in Neural Information Processing Systems 32
2019
Later among the works it cites.
Twitter, “Number of monthly active twitter users worldwide from 1st quarter 2010 to 1st quarter 2019 (in millions),” 2019
2019
Later among the works it cites.
C. Sanderson and R. Curtin, “Practical sparse matrices in c++ with hybrid storage and template-based expression optimisation,” Mathematical and Computational Applications
2019
Later among the works it cites.
K. Hegde, H. Asghari-Moghaddam, M. Pellauer, N. Crago, A. Jaleel, E. Solomonik, J. Emer, and C. W. Fletcher, “Extensor: An accelerator for sparse tensor algebra,” in MICRO ’52
2019
Later among the works it cites.
A. Gondimalla, N. Chesnut, M. Thottethodi, and T. N. Vijaykumar, “Sparten: A sparse tensor accelerator for convolutional neural networks,” MICRO ’52, pp. 151–165, 2019
2019
Later among the works it cites.
M. Zhu, T. Zhang, Z. Gu, and Y. Xie, “Sparse tensor core: Algorithm and hardware co-design for vector-wise sparse neural networks on modern gpus,” MICRO ’52, pp. 359–371, ACM, 2019
2019
Later among the works it cites.
H. Wang, K. Wang, J. Yang, N. Sun, H.-S. Lee, and S. Han, “Tts: Transferable transistor sizing with graph neural networks and reinforcement learning,” DAC 57, 2020
2020
Closest in time.