Fetching the paper…
Reading the bibliography…
Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition.
Multi–level adaptive solutions to boundary–value problems
A. Brandt · 1977
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. E. Boser, and J. S. Denker · 1990
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A multigrid tutorial
W. L. Briggs, V. E. Henson, and S. F. McCormick · 2000
Earlier work this paper cites.
An introduction to the adjoint approach to design
M. B. Giles and N. A. Pierce · 2000
Earlier work this paper cites.
Multigrid
U. Trottenberg, C. Oosterlee, and A. Schüller · 2001
Earlier work this paper cites.
Parallel Lagrange–Newton–Krylov–Schur Methods for PDE-Constrained Optimization. Part I: The Krylov–Schur Solver
G. Biros and O. Ghattas · 2005
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
Analysis of the parareal time-parallel time-integration method
M. J. Gander and S. Vandewalle · 2007
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation
A. Griewank and A. Walther · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Y. Bengio et al · 2009
Earlier work this paper cites.
Single-step one-shot aerodynamic shape optimization
N. Gauger and E. Özkaya · 2009
Earlier work this paper cites.
Minimal repetition dynamic checkpointing algorithm for unsteady adjoint calculation
Q. Wang, P. Moin, and G. Iaccarino · 2009
Earlier work this paper cites.
Approximate nullspace iterations for KKT systems
K. Ito, K. Kunisch, V. Schulz, and I. Gherman · 2010
Earlier work this paper cites.
Computational optimization of systems governed by partial differential equations
A. Borzì and V. Schulz · 2011
Earlier work this paper cites.
Natural Language Processing (Almost) from Scratch
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa · 2011
Earlier work this paper cites.
Adaptive multilevel inexact SQP methods for PDE-constrained optimization
J. C. Ziems and S. Ulbrich · 2011
Earlier work this paper cites.
Learning from data
Y. S. Abu-Mostafa, M. Magdon-Ismail, and H.-T. Lin · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Question Answering with Subgraph Embeddings
A. Bordes, S. Chopra, and J. Weston · 2014
Cited alongside, same era.
Optimal design with bounded retardation for problems with non-separable adjoints
T. Bosse, N. Gauger, A. Griewank, S. Günther, L. Kaland, and et al · 2014
Cited alongside, same era.
Distributed training of deep neural networks: Theoretical and practical limits of parallel scalability
J. Keuper and F.-J. Preundt · 2016
Later among the works it cites.
Two-level convergence theory for multigrid reduction in time (MGRIT)
V. Dobrev, T. Kolev, N. Petersson, and J. Schroder · 2017
Later among the works it cites.
A Proposal on Machine Learning via Dynamical Systems
W. E · 2017
Later among the works it cites.
Multigrid methods with space–time concurrency
R. D. Falgout, S. Friedhoff, T. V. Kolev, S. P. MacLachlan, J. B. Schroder, and S. Vandewalle · 2017
Later among the works it cites.
Multigrid reduction in time for nonlinear parabolic problems: A case study
R. D. Falgout, T. A. Manteuffel, B. O’Neill, and J. B. Schroder · 2017
Later among the works it cites.
The reversible residual network: Backpropagation without storing activations
A. N. Gomez, M. Ren, R. Urtasun, and R. B. Grosse · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One-shot approaches to design optimzation
T. Bosse, N. Gauger, A. Griewank, S. Günther, and V. Schulz · 2014
Cited alongside, same era.
Adaptive sequencing of primal, dual, and design steps in simulation based optimization
T. Bosse, L. Lehmann, and A. Griewank · 2014
Cited alongside, same era.
Parallel time integration with multigrid
R. D. Falgout, S. Friedhoff, T. V. Kolev, S. P. MacLachlan, and J. B. Schroder · 2014
Cited alongside, same era.
On Using Very Large Target Vocabulary for Neural Machine Translation
S. Jean, K. Cho, R. Memisevic, and Y. Bengio · 2014
Cited alongside, same era.
220 band aviris hyperspectral image data set: June 12, 1992 indian pine test site 3, Sep 2015
M. F. Baumgardner, L. L. Biehl, and D. A. Landgrebe · 2015
Cited alongside, same era.
50 years of time parallel time integration
M. J. Gander · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Later among the works it cites.
Stable architectures for deep neural networks
E. Haber and L. Ruthotto · 2017
Later among the works it cites.
Y. Lu, A. Zhong, Q. Li, and B. Dong · 2017
Later among the works it cites.
Parallelizing over artificial neural network training runs with multigrid
J. B. Schroder · 2017
Later among the works it cites.
Reversible architectures for arbitrarily deep residual neural networks
B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham · 2018
Closest in time.
Neural ordinary differential equations
T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud · 2018
Closest in time.
A nonlinear ParaExp algorithm
M. J. Gander, S. Güttel, and M. Petcu · 2018
Closest in time.
A non-intrusive parallel-in-time adjoint solver with the XBraid library
S. Günther, N. R. Gauger, and J. B. Schroder · 2018
Closest in time.
A non-intrusive parallel-in-time approach for simultaneous optimization with unsteady pdes
S. Günther, N. R. Gauger, and J. B. Schroder · 2018
Closest in time.
Learning across scales - A multiscale method for convolution neural networks
E. Haber, L. Ruthotto, E. Holtham, and S.-H. Jun · 2018
Closest in time.
Pipedream: Fast and efficient pipeline parallel dnn training, 2018
A. Harlap, D. Narayanan, A. Phanishayee, V. Seshadri, N. Devanur, G. Ganger, and P. Gibbons · 2018
Closest in time.
An efficient parallel-in-time method for optimization with parabolic pdes
S. Götschel and M. L. Minion · 2019
Closest in time.