Fetching the paper…
Reading the bibliography…
On many parallel machines, the time LQCD applications spent in communication is a significant contribution to the total wall-clock time, especially in the strong-scaling limit.
2014
Earlier work this paper cites.
P. Arts et al., QPACE 2 and Domain Decomposition on the Intel Xeon Phi
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
https://github.com/pjgeorg/pMR
Cited in the paper.
S. Heybrock et al., Adaptive algebraic multigrid on SIMD architectures
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…