Fetching the paper…
Reading the bibliography…
Tensor computations--in particular tensor contraction (TC)--are important kernels in many scientific computing applications.
Basic linear algebra subprograms for FORTRAN usage
C. L. Lawson, R. J. Hanson, D. R. Kincaid, and F. T. Krogh · 1979
Earlier work this paper cites.
An extended set of FORTRAN basic linear algebra subprograms
J. J. Dongarra, J. Du Croz, S. Hammarling, and R. J. Hanson · 1988
Earlier work this paper cites.
A fifth-order perturbation comparison of electron correlation theories
K. Raghavachari, G. W. Trucks, J. A. Pople, and M. Head-Gordon · 1989
Earlier work this paper cites.
A set of level 3 basic linear algebra subprograms
J. J. Dongarra, J. Du Croz, S. Hammarling, and I. S. Duff · 1990
Earlier work this paper cites.
A direct product decomposition approach for symmetry exploitation in many-body methods. I. Energy calculations
J. F. Stanton, J. Gauss, J. D. Watts, and R. J. Bartlett · 1991
Earlier work this paper cites.
Arrays in Blitz++
T. L. Veldhuizen · 1998
Earlier work this paper cites.
Multilinear analysis of image ensembles: TensorFaces
M. A. O. Vasilescu and D. Terzopoulos · 2002
Earlier work this paper cites.
Tensor contraction engine: Abstraction and automated parallel implementation of configuration-interaction, coupled-cluster, and many-body perturbation theories
S. Hirata · 2003
Earlier work this paper cites.
Multi-way analysis: Applications in the chemical sciences
A. Smilde, R. Bro, and P. Geladi · 2005
Earlier work this paper cites.
Algorithm 862: MATLAB tensor classes for fast algorithm prototyping
B. W. Bader and T. G. Kolda · 2006
Earlier work this paper cites.
Coupled-cluster theory in quantum chemistry
R. J. Bartlett and M. Musiał · 2007
Earlier work this paper cites.
Anatomy of high-performance matrix multiplication
K. Goto and R. A. van de Geijn · 2008
Earlier work this paper cites.
High-performance implementation of the level-3 BLAS
K. Goto and R. A. van de Geijn · 2008
Earlier work this paper cites.
Applied multiway data analysis
P. M. Kroonenberg · 2008
Earlier work this paper cites.
Automating the generation of composed linear algebra kernels
G. Belter, E. R. Jessup, I. Karlin, and J. G. Siek · 2009
Cited alongside, same era.
Performance optimization of tensor contraction expressions for many-body methods in quantum chemistry
A. Hartono, Q. Lu, T. Henretty, S. Krishnamoorthy, H. Zhang, G. Baumgartner, D. E. Bernholdt, M. Nooijen, R. Pitzer, J. Ramanujam, and P. Sadayappan · 2009
Cited alongside, same era.
Tensor decompositions and applications
T. Kolda and B. Bader · 2009
Cited alongside, same era.
Eigen v3, 2010
G. Guennebaud, B. Jacob, et al · 2010
Cited alongside, same era.
An efficient matrix-matrix multiplication based antisymmetric tensor contraction engine for general order coupled cluster
M. Hanrath and A. Engels-Putzka · 2010
Cited alongside, same era.
Optimizing tensor contraction expressions for hybrid CPU-GPU execution
W. Ma, S. Krishnamoorthy, O. Villa, K. Kowalski, and G. Agrawal · 2011
On the performance prediction of BLAS-based tensor contractions
E. Peise, D. Fabregat-Traver, and P. Bientinesi · 2014
Later among the works it cites.
Anatomy of high-performance many-threaded matrix multiplication
T. M. Smith, R. A. van de Geijn, M. Smelyanskiy, J. R. Hammond, and F. G. Van Zee · 2014
Later among the works it cites.
A massively parallel tensor contraction framework for coupled-cluster computations
E. Solomonik, D. Matthews, J. R. Hammond, J. F. Stanton, and J. Demmel · 2014
Later among the works it cites.
Scalable task-based algorithm for multiplication of block-rank-sparse matrices
J. A. Calvin, C. A. Lewis, and E. F. Valeev · 2015
Later among the works it cites.
Task-based algorithm for matrix multiplication: A step towards block-sparse tensor computing
J. A. Calvin and E. F. Valeev · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The NumPy array: A structure for efficient numerical computation
S. van de Walt, S. C. Colbert, and G. Varoquaux · 2011
Cited alongside, same era.
Model-driven level 3 BLAS performance optimization on Loongson 3a processor
X. Zhang, Q. Wang, and Y. Zhang · 2012
Cited alongside, same era.
New implementation of high-level correlated methods using a general block-tensor library for high-performance electronic structure calculations
E. Epifanovsky, M. Wormit, T. Kus, A. Landau, D. Zuev, K. Khistyaev, P. Manohar, I. Kaliman, A. Dreuw, and A. I. Krylov · 2013
Cited alongside, same era.
A case study in mechanically deriving dense linear algebra code
B. Marker, D. Batory, and R. A. van de Geijn · 2013
Cited alongside, same era.
AUGEM: Automatically generate high performance dense linear algebra kernels on x86 CPUs
Q. Wang, X. Zhang, Y. Zhang, and Q. Yi · 2013
Cited alongside, same era.
Efficient Implementation of Many-body Quantum Chemical Methods on the Intel® Xeon Phi™ Coprocessor
E. Aprà , M. Klemm, and K. Kowalski · 2014
Cited alongside, same era.
An input-adaptive and in-place approach to dense tensor-times-matrix multiply
J. Li, C. Battaglino, I. Perros, J. Sun, and R. Vuduc · 2015
Later among the works it cites.
An efficient tensor transpose algorithm for multicore CPU, Intel Xeon Phi, and NVidia Tesla GPU
Dmitry I. Lyakh · 2015
Later among the works it cites.
Non-orthogonal spin-adaptation of coupled cluster methods: A new implementation of methods including quadruple excitations
D. A. Matthews and J. F. Stanton · 2015
Later among the works it cites.
BLIS: A framework for rapidly instantiating BLAS functionality
F. G. Van Zee and R. A. van de Geijn · 2015
Later among the works it cites.
Performance optimization for the K-nearest neighbors kernel on x86 architectures
C. D. Yu, J. Huang, W. Austin, B. Xiao, and G. Biros · 2015
Later among the works it cites.
Strassen’s algorithm reloaded
J. Huang, T. M. Smith, G. M. Henry, and R. A. van de Geijn · 2016
Closest in time.
Design of a high-performance GEMM-like tensor-tensor multiplication
P. Springer and P. Bientinesi · 2016
Closest in time.
TTC: A high-performance compiler for tensor transpositions
P. Springer, J. R. Hammond, and P. Bientinesi · 2016
Closest in time.