Fetching the paper…
Reading the bibliography…
Conventional GPU implementations of Strassen's algorithm (Strassen) typically rely on the existing high-performance matrix multiplication (GEMM), trading space for time.
Gaussian elimination is not optimal
Volker Strassen · 1969
Earlier work this paper cites.
A set of level 3 basic linear algebra subprograms
Jack J. Dongarra, Jeremy Du Croz, Sven Hammarling, and Iain Duff · 1990
Earlier work this paper cites.
Implementation of Strassen’s algorithm for matrix multiplication
Steven Huss-Lederman, Elaine M. Jacobson, Anna Tsao, Thomas Turnbull, and Jeremy R. Johnson · 1996
Earlier work this paper cites.
Tuning Strassen’s matrix multiplication for memory efficiency
Mithuna Thottethodi, Siddhartha Chatterjee, and Alvin R. Lebeck · 1998
Earlier work this paper cites.
Accuracy and Stability of Numerical Algorithms
Nicholas J. Higham · 2002
Earlier work this paper cites.
Fast matrix multiplication is stable
James Demmel, Ioana Dumitriu, Olga Holtz, and Robert Kleinberg · 2007
Earlier work this paper cites.
Benchmarking GPUs to tune dense linear algebra
Vasily Volkov and James W Demmel · 2008
Earlier work this paper cites.
An improved magma gemm for fermi graphics processing units
Rajib Nath, Stanimire Tomov, and Jack Dongarra · 2010
Earlier work this paper cites.
Exploiting parallelism in matrix-computation kernels for symmetric multiprocessor systems: Matrix-multiplication and matrix-addition algorithm optimizations by software pipelining and threads allocation
Paolo D’Alberto, Marco Bodrato, and Alexandru Nicolau · 2011
Earlier work this paper cites.
Strassen’s matrix multiplication on GPUs
Junjie Li, Sanjay Ranka, and Sartaj Sahni · 2011
Earlier work this paper cites.
Fast implementation of dgemm on fermi gpu
Guangming Tan, Linchuan Li, Sean Triechle, Everett Phillips, Yungang Bao, and Ninghui Sun · 2011
Earlier work this paper cites.
Combining in-situ and in-transit processing to enable extreme-scale scientific analysis
Janine C Bennett, Hasan Abbasi, Peer-Timo Bremer, Ray Grout, Attila Gyulassy, Tong Jin, Scott Klasky, Hemanth Kolla, Manish Parashar, Valerio Pascucci, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Performance upper bound analysis and optimization of sgemm on fermi and kepler gpus
Junjie Lai and Andre Seznec · 2013
Cited alongside, same era.
Accelerating Strassen-Winograd’s matrix multiplication algorithm on GPUs
Pai-Wei Lai, Humayun Arafat, Venmugil Elango, and P. Sadayappan · 2013
Cited alongside, same era.
A processing in memory taxonomy and a case for studying fixed-function pim
Gabriel H Loh, Nuwan Jayasena, M Oskin, Mark Nutter, David Roberts, Mitesh Meswani, Dong Ping Zhang, and Mike Ignatowski · 2013
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Generating families of practical fast matrix multiplication algorithms
Jianyu Huang, Leslie Rice, Devin A. Matthews, and Robert A. van de Geijn · 2017
Later among the works it cites.
CUTLASS: Fast linear algebra in CUDA C++
Andrew Kerr, Duane Merrill, Julien Demouth, and John Tran · 2017
Later among the works it cites.
Understanding the gpu microarchitecture to achieve bare-metal performance tuning
Xiuxia Zhang, Guangming Tan, Shuangbai Xue, Jiajia Li, Keren Zhou, and Mingyu Chen · 2017
Later among the works it cites.
GitHub Repository, 2018
CUTLASS: CUDA templates for linear algebra subroutines (v0.1.0) · 2018
Closest in time.
Strassen’s algorithm for tensor contraction
Jianyu Huang, Devin A. Matthews, and Robert A. van de Geijn · 2018
Closest in time.
Dissecting the nvidia volta gpu architecture via microbenchmarking
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pim-enabled instructions: a low-overhead, locality-aware processing-in-memory architecture
Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi · 2015
Cited alongside, same era.
A framework for practical parallel fast matrix multiplication
Austin R. Benson and Grey Ballard · 2015
Cited alongside, same era.
Performance optimization for the K-Nearest Neighbors kernel on x86 architectures
Chenhan D. Yu, Jianyu Huang, Woody Austin, Bo Xiao, and George Biros · 2015
Cited alongside, same era.
Improving the numerical stability of fast matrix multiplication
Grey Ballard, Austin R Benson, Alex Druinsky, Benjamin Lipshitz, and Oded Schwartz · 2016
Cited alongside, same era.
Strassen’s algorithm reloaded
Jianyu Huang, Tyler M. Smith, Greg M. Henry, and Robert A. van de Geijn · 2016
Cited alongside, same era.
Inside Volta: The world’s most advanced data center GPU
Luke Durant, Olivier Giroux, Mark Harris, and Nick Stam · 2017
Cited alongside, same era.
Nervanagpu
Scott Gray · 2017
Cited alongside, same era.
Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele P Scarpazza · 2018
Closest in time.
High-performance tensor contraction without transposition
Devin A. Matthews · 2018
Closest in time.
CUDA C programming guide
NVIDIA · 2018
Closest in time.
CUDA occupancy calculator
NVIDIA · 2018
Closest in time.
Nvidia cuBLAS
NVIDIA · 2018
Closest in time.
Tuning CUDA applications for volta
NVIDIA · 2018
Closest in time.
Design of a high-performance GEMM-like tensor-tensor multiplication
Paul Springer and Paolo Bientinesi · 2018
Closest in time.