Fetching the paper…
Reading the bibliography…
A Multigrid Full Approximation Storage algorithm for solving Deep Residual Networks is developed to enable neural network parallelized layer-wise training and concurrent computational kernel execution on GPUs.
D. J. Mavriplis, and A. Jameson, “Multigrid solution of the Navier-Stokes equations on triangular meshes,” AIAA journal 28, no. 8 (1990): 1415-1425
1990
Earlier work this paper cites.
D. J. Mavriplis, “Multigrid Techniques for Unstructured Mesh,” Numerical Methods for Fluid Dynamics V 5 (1995): 129
1995
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE 86, no. 11 (1998): 2278-2324
1998
Earlier work this paper cites.
W. Gropp, E. Lusk, A. Skjellum, Using MPI: portable parallel programming with the message-passing interface
1999
Earlier work this paper cites.
J. Nickolls, I. Buck, M. Garland, and K. Skadron, “Scalable parallel programming with CUDA,” Queue 6, no. 2 (2008): 40-53
2008
Earlier work this paper cites.
2014
Earlier work this paper cites.
R. D. Falgout, S. Friedhoff, T. V. Kolev, S. P. MacLachlan, and J. B. Schroder. “Parallel time integration with multigrid,” SIAM J. Sci. Comput., 36(6):C635–C661, 2014. LLNL-JRNL645325
2014
Earlier work this paper cites.
R. D. Falgout, S. Friedhoff, T. V. Kolev, S. P. MacLachlan, and J. B. Schroder. ”Parallel time integration with multigrid,” SIAM J. Sci. Comput., 36(6):C635–C661, 2014. LLNL-JRNL645325
2014
Earlier work this paper cites.
T. Bosse, , N. R. Gauger, A. Griewank, S. Günther, and V. Schulz, “One-shot approaches to design optimzation,” In Trends in PDE Constrained Optimization, pp. 43-66. Birkhäuser, Cham, 2014
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and Li Fei-Fei, “Imagenet large scale visual recognition challenge,” International journal of computer vision 115, no. 3 (2015): 211-252
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778. 2016
2016
Cited alongside, same era.
Y. Lu, A. Zhong, Q. Li, and B. Dong, “Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations,” arXiv preprint arXiv: 1710.10121 (2017)
2017
Cited alongside, same era.
V. A. Dobrev, T. Kolev, N. A. Petersson, and J. B. Schroder, “Two-level convergence theory for multigrid reduction in time (MGRIT),” SIAM Journal on Scientific Computing 39, no. 5 (2017): S501-S527
2017
Cited alongside, same era.
T. Ben-Nun and T. Hoefler, “Demystifying parallel and distributed deep learning: An in-depth concurrency analysis,” ACM Computing Surveys (CSUR) 52, no. 4 (2019): 1-43
2019
Later among the works it cites.
2019
Later among the works it cites.
V. Gadepally, J. Goodwin, J. Kepner, A. Reuther, H. Reynolds, S. Samsi, J. Su, and D. Martinez, “AI Enabling Technologies: A Survey,” arXiv preprint arXiv: 1905.03592 (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. D. Falgout, S. Friedhoff, T. V. Kolev, S. P. MacLachlan, J. B. Schroder, and S. Vandewalle, “Multigrid methods with space–time concurrency,” Computing and Visualization in Science, 18(4-5):123–143, 2017
2017
Cited alongside, same era.
P. Messina, “The exascale computing project,” Computing in Science & Engineering 19, no. 3 (2017): 63-67
2017
Cited alongside, same era.
D. Amodei, and D. Hernandez, “AI and Compute,” https://openai.com/blog/ai-and-compute/
2018
Cited alongside, same era.
T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud, “Neural ordinary differential equations,” In Advances in neural information processing systems, pp. 6571-6583. 2018
2018
Cited alongside, same era.
A. Reuther, J. Kepner, C. Byun, S. Samsi, W. Arcand, D. Bestor, B. Bergeron et
2018
Cited alongside, same era.
S. Günther, N. R. Gauger, and J. B. Schroder, “A non-intrusive parallel-in-time approach for simultaneous optimization with unsteady PDEs,” Optimization Methods and Software 34, no. 6 (2019): 1306-1321
2019
Later among the works it cites.
2019
Later among the works it cites.
C. E. Leiserson, N. C. Thompson, J. S. Emer, B. C. Kuszmaul, B. W. Lampson, D. Sanchez, and T. B. Schardl, “There’s plenty of room at the Top: What will drive computer performance after Moore’s law?,” Science 368.6495, June 2020. DOI: 10.1126/science.aam9744
2020
Closest in time.
L. Gaedke-Merzhäuser, A. Kopanic̆áková, and R. Krause, “Multilevel Minimization for Deep Residual Networks,” arXiv preprint arXiv: 2004.06196 (2020)
2020
Closest in time.
S. G u ¨ \ddot{u} nther, L. Ruthotto, J. B. Schroder, E. C. Cyr, and N. R. Gauger, “Layer-parallel training of deep residual neural networks,” SIAM Journal on Mathematics of Data Science 2, no. 1 (2020): 1-23
2020
Closest in time.