Fetching the paper…
Reading the bibliography…
Gradient-based algorithms for training ResNets typically require a forward pass of the input data, followed by back-propagating the objective gradient to update parameters, which are time-consuming for deep ResNets.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1985
Earlier work this paper cites.
The calculation of posterior distributions by data augmentation
M. A. Tanner and W. H. Wong · 1987
Earlier work this paper cites.
Reevaluating amdahl’s law
J. L. Gustafson · 1988
Earlier work this paper cites.
Theory of the backpropagation neural network
R. Hecht-Nielsen · 1992
Earlier work this paper cites.
A parareal in time procedure for the control of partial differential equations
Y. Maday and G. Turinici · 2002
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
Optimal control of partial differential equations: theory, methods, and applications , volume 112
F. Tröltzsch · 2010
Earlier work this paper cites.
Calculus of variations and optimal control theory: a concise introduction
D. Liberzon · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, et al · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Gpu asynchronous stochastic gradient descent to speed up neural network training
T. Paine, H. Jin, J. Yang, Z. Lin, and T. Huang · 2013
Earlier work this paper cites.
Distributed optimization of deeply nested systems
M. Carreira-Perpinan and W. Wang · 2014
Earlier work this paper cites.
Multiple Shooting and Time Domain Decomposition Methods
T. Carraro, M. Geiger, S. Rorkel, and R. Rannacher · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
F. N. Iandola, M. W. Moskewicz, K. Ashraf, and K. Keutzer · 2016
Cited alongside, same era.
Training neural networks without gradients: A scalable admm approach
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, and T. Goldstein · 2016
Cited alongside, same era.
A proposal on machine learning via dynamical systems
W. E · 2017
Cited alongside, same era.
Decoupled neural interfaces using synthetic gradients
M. Jaderberg, W. M. Czarnecki, S. Osindero, O. Vinyals, A. Graves, D. Silver, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Learning across scales—multiscale methods for convolution neural networks
E. Haber, L. Ruthotto, E. Holtham, and S.-H. Jun · 2018
Later among the works it cites.
Pipedream: Fast and efficient pipeline parallel dnn training
A. Harlap, D. Narayanan, A. Phanishayee, V. Seshadri, N. Devanur, G. Ganger, and P. Gibbons · 2018
Later among the works it cites.
Decoupled parallel backpropagation with convergence guarantee
Z. Huo, B. Gu, Q. Yang, and H. Huang · 2018
Later among the works it cites.
Deep limits of residual neural networks
M. Thorpe and Y. van Gennip · 2018
Later among the works it cites.
Global convergence in deep learning with variable splitting via the kurdyka-łojasiewicz property
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maximum principle based algorithms for deep learning
Q. Li, L. Chen, C. Tai, and E. Weinan · 2017
Cited alongside, same era.
Y. Lu, A. Zhong, Q. Li, and B. Dong · 2017
Cited alongside, same era.
Synthetic gradient methods with virtual forward-backward networks
T. Miyato, D. Okanohara, S.-i. Maeda, and M. Koyama · 2017
Cited alongside, same era.
Neural ordinary differential equations
R. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud · 2018
Cited alongside, same era.
Beyond backprop: Online alternating minimization with auxiliary variables
A. Choromanska, B. Cowen, S. Kumaravel, R. Luss, M. Rigotti, I. Rish, B. Kingsbury, P. DiAchille, V. Gurev, R. Tejwani, et al · 2018
Cited alongside, same era.
Decoupling backpropagation using constrained optimization methods
A. Gotmare, V. Thomas, J. Brea, and M. Jaggi · 2018
Cited alongside, same era.
J. Zeng, S. Ouyang, T. T.-K. Lau, S. Lin, and Y. Yao · 2018
Later among the works it cites.
Anode: Unconditionally accurate memory-efficient gradients for neural odes
A. Gholami, K. Keutzer, and G. Biros · 2019
Later among the works it cites.
Predict globally, correct locally: Parallel-in-time optimal control of neural networks
P. Parpas and C. Muir · 2019
Later among the works it cites.
A convergence analysis of nonlinearly constrained admm in deep learning
J. Zeng, S.-B. Lin, and Y. Yao · 2019
Later among the works it cites.
Layer-parallel training of deep residual neural networks
S. G u ¨ \ddot{\textnormal{u}} nther, L. Ruthotto, J. B. Schroder, E. C. Cyr, and N. R. Gauger · 2020
Closest in time.
Layer-parallel training with gpu concurrency of deep residual neural networks via nonlinear multigrid
A. Kirby, S. Samsi, M. Jones, A. Reuther, J. Kepner, and V. Gadepally · 2020
Closest in time.
Training neural networks by lifted proximal operator machines
J. Li, M. Xiao, C. Fang, Y. Dai, C. Xu, and Z. Lin · 2020
Closest in time.
Local propagation in constraint-based neural network
G. Marra, M. Tiezzi, S. Melacci, A. Betti, M. Maggini, and M. Gori · 2020
Closest in time.