Fetching the paper…
Reading the bibliography…
Modern graphics hardware is designed for highly parallel numerical tasks and promises significant cost and performance benefits for many scientific applications.
W. Kahan, “Further remarks on reducing truncation errors,” Communications of the ACM
1965
Earlier work this paper cites.
R. S. Martin, G. Peters and J. H. Wilkinson, “Handbook series linear algebra: Iterative refinement of the solution of a positive definite system of equations,” Numerische Mathematik
1966
Earlier work this paper cites.
P. De Forcrand, D. Lellouch and C. Roiesnel, “Optimizing a lattice QCD simulation program,” J. Comput. Phys
1985
Earlier work this paper cites.
1986
Earlier work this paper cites.
T. A. DeGrand and P. Rossi, Comput. Phys. Commun
1990
Earlier work this paper cites.
G. L. G. Sleijpen, and H. A. van der Vorst, “Reliable updated residuals in hybrid Bi-CG methods,” Computing
1996
Earlier work this paper cites.
D. Göddeke, R. Strzodka and S. Turek, “Accelerating Double Precision FEM Simulations with GPUs” Proceedings of ASIM 2005 - 18th Symposium on Simulation Technique (2005)
2005
Earlier work this paper cites.
R. G. Edwards and B. Joo [SciDAC Collaboration and LHPC Collaboration and UKQCD Collaboration], “The Chroma software system for lattice QCD,” Nucl. Phys. Proc. Suppl. 140
2005
Cited alongside, same era.
R. Strzodka and D. Göddeke, “Pipelined Mixed Precision Algorithms on FPGAs for Fast and Accurate PDE Solvers from Low Precision Components,” IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM 2006)
2006
Cited alongside, same era.
G. I. Egri, Z. Fodor, C. Hoelbling, S. D. Katz, D. Nogradi and K. K. Szabo, “Lattice QCD as a video game,” Comput. Phys. Commun. 177
2007
Cited alongside, same era.
M. Harris, “Optimizing parallel reduction in CUDA,” presentation packaged with CUDA Toolkit, NVIDIA Corporation (2007)
2007
Cited alongside, same era.
2008
Later among the works it cites.
NVIDIA Corporation, “NVIDIA CUDA Programming Guide” (2009), http://developer.download.nvidia.com/compute/cuda/2_3/toolkit/docs/NVIDIA_CUDA_Programming_Guide_2.3.pdf
2009
Closest in time.
A. Munshi et al., “The OpenCL specification version 1.0,” Technical report, Khronos OpenCL Working Group, 2009.04.02 (2009)
2009
Closest in time.
G. Ruetsch and P. Micikevicius, “Optimizing matrix transpose in CUDA,” NVIDIA Technical Report
2009
Closest in time.
D. Holmgren, “Fermilab Status” (2009), http://www.usqcd.org/meetings/allHands2009/slides/holmgren_allhands_2009.pdf
2009
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2008
Cited alongside, same era.
N. Bell and M. Garland, “Efficient sparse matrix-vector multiplication on CUDA,” NVIDIA Technical Report
2008
Cited alongside, same era.
2008
Cited alongside, same era.
http://www.python.org
Cited in the paper.
http://lattice.bu.edu/quda
Cited in the paper.
http://usqcd.jlab.org/usqcd-docs/chroma
Cited in the paper.
http://qcdoc.phys.columbia.edu/cps.html
Cited in the paper.
http://usqcd.jlab.org/usqcd-docs/qdp
Cited in the paper.
2009
Closest in time.