Fetching the paper…
Reading the bibliography…
The training of deep residual neural networks (ResNets) with backpropagation has a memory cost that increases linearly with respect to the depth of the network.
The theory of matrices
Gantmacher, F. R · 1959
Earlier work this paper cites.
On the existence and uniqueness of the real logarithm of a matrix
Culver, W. J · 1966
Earlier work this paper cites.
Zeros of entire functions
Runckel, H.-J · 1969
Earlier work this paper cites.
Computational graphs and rounding error
Bauer, F. L · 1974
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Dense sets of diagonalizable matrices
Hartfiel, D. J · 1995
Earlier work this paper cites.
Lectures on entire functions , volume 150
Levin, B. Y · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Approximation by entire functions belonging to the laguerre–polya class
Dryanov, D. and Rahman, Q · 1999
Earlier work this paper cites.
An introduction to automatic differentiation
Verma, A · 2000
Earlier work this paper cites.
Iterated laguerre and turán inequalities
Craven, T. and Csordas, G · 2002
Earlier work this paper cites.
An iterative thresholding algorithm for linear inverse problems with a sparsity constraint
Daubechies, I., Defrise, M., and De Mol, C · 2004
Earlier work this paper cites.
Geometric numerical integration: structure-preserving algorithms for ordinary differential equations , volume 31
Hairer, E., Lubich, C., and Wanner, G · 2006
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Griewank, A. and Walther, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
The image of the exponential map and some applications
Andrica, D. and Rohan, R.-A · 2010
Earlier work this paper cites.
Learning fast approximations of sparse coding
Gregor, K. and LeCun, Y · 2010
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Krizhevsky, A., Nair, V., and Hinton, G · 2010
Earlier work this paper cites.
Functions of one complex variable II , volume 159
Conway, J. B · 2012
Earlier work this paper cites.
Training deep and recurrent networks with hessian-free optimization
Martens, J. and Sutskever, I · 2012
Earlier work this paper cites.
Differential equations and dynamical systems , volume 7
Perko, L · 2013
Cited alongside, same era.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., et al · 2016
Cited alongside, same era.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Stochastic training of residual networks: a differential equation viewpoint
Sun, Q., Tao, Y., and Du, Q · 2018
Later among the works it cites.
Superneurons: Dynamic gpu memory management for training deep neural networks
Wang, L., Ye, J., Zhao, Y., Wu, W., Li, A., Song, S. L., Xu, Z., and Kraska, T · 2018
Later among the works it cites.
Invertible residual networks
Behrmann, J., Grathwohl, W., Chen, R. T., Duvenaud, D., and Jacobsen, J.-H · 2019
Later among the works it cites.
Anode: Unconditionally accurate memory-efficient gradients for neural odes
Gholami, A., Keutzer, K., and Biros, G · 2019
Later among the works it cites.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Cited alongside, same era.
Convolutional neural networks for medical image analysis: Full training or fine tuning?
Tajbakhsh, N., Shin, J. Y., Gurudu, S. R., Hurst, R. T., Kendall, C. B., Gotway, M. B., and Liang, J · 2016
Cited alongside, same era.
The reversible residual network: Backpropagation without storing activations
Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B · 2017
Cited alongside, same era.
Stable architectures for deep neural networks
Haber, E. and Ruthotto, L · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Large kernel matters–improve semantic segmentation by global convolutional network
Peng, C., Zhang, X., Yu, G., Luo, G., and Sun, J · 2017
Cited alongside, same era.
Later among the works it cites.
Deep learning via dynamical systems: An approximation perspective
Li, Q., Lin, T., and Shen, Z · 2019
Later among the works it cites.
Deep neural networks motivated by partial differential equations
Ruthotto, L. and Haber, E · 2019
Later among the works it cites.
Augmented neural odes
Teh, Y., Doucet, A., and Dupont, E · 2019
Later among the works it cites.
In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019
Touvron, H., Vedaldi, A., Douze, M., and Jegou, H · 2019
Later among the works it cites.
A mean-field optimal control formulation of deep learning
Weinan, E., Han, J., and Li, Q · 2019
Later among the works it cites.
Anodev2: A coupled neural ode evolution framework
Zhang, T., Yao, Z., Gholami, A., Keutzer, K., Gonzalez, J., Biros, G., and Mahoney, M · 2019
Later among the works it cites.
Momentum-net: Fast and convergent iterative neural network for inverse problems
Chun, I. Y., Huang, Z., Lim, H., and Fessler, J · 2020
Later among the works it cites.
Towards understanding normalization in neural odes
Gusak, J., Markeeva, L., Daulbaev, T., Katrutsa, A., Cichocki, A., and Oseledets, I · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Later among the works it cites.
Dissecting neural odes
Massaroli, S., Poli, M., Park, J., Yamashita, A., and Asama, H · 2020
Later among the works it cites.
Momentumrnn: Integrating momentum into recurrent neural networks
Nguyen, T. M., Baraniuk, R. G., Bertozzi, A. L., Osher, S. J., and Wang, B · 2020
Later among the works it cites.
On second order behaviour in augmented neural odes
Norcliffe, A., Bodnar, C., Day, B., Simidjievski, N., and Liò, P · 2020
Later among the works it cites.
Continuous-in-depth neural networks
Queiruga, A. F., Erichson, N. B., Taylor, D., and Mahoney, M. W · 2020
Later among the works it cites.
Universal approximation property of neural ordinary differential equations
Teshima, T., Tojo, K., Ikeda, M., Ishikawa, I., and Oono, K · 2020
Later among the works it cites.
Approximation capabilities of neural odes and invertible residual networks
Zhang, H., Gao, X., Unterman, J., and Arodz, T · 2020
Later among the works it cites.
Coupled oscillatory recurrent neural network (co{rnn}): An accurate and (gradient) stable architecture for learning long time dependencies
Rusch, T. K. and Mishra, S · 2021
Closest in time.