Fetching the paper…
Reading the bibliography…
A deep equilibrium model uses implicit layers, which are implicitly defined through an equilibrium point of an infinite sequence of computation.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Zur theorie der matrices
Oskar Perron · 1907
Earlier work this paper cites.
Über matrizen aus nicht negativen elementen
Georg Frobenius · 1912
Earlier work this paper cites.
Convergence properties of the spline fit
J Harold Ahlberg and Edwin N Nilson · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
A lower bound for the smallest singular value of a matrix
James M Varah · 1975
Earlier work this paper cites.
Reverse accumulation and attractive fixed points
Bruce Christianson · 1994
Earlier work this paper cites.
Semeion handwritten digit data set
B Tactile Srl and Italy Brescia · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Matrix differentiation
Randal J Barnes · 2006
Earlier work this paper cites.
Algorithmic differentiation of implicit functions and optimal values
Bradley M Bell and James V Burke · 2008
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Andreas Griewank and Andrea Walther · 2008
Earlier work this paper cites.
Bounds for norms of the matrix inverse and the smallest singular value
Nenad Morača · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Cited alongside, same era.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Cited alongside, same era.
Theory of deep learning iii: explaining the non-overfitting puzzle
Tomaso Poggio, Kenji Kawaguchi, Qianli Liao, Brando Miranda, Lorenzo Rosasco, Xavier Boix, Jack Hidary, and Hrushikesh Mhaskar · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Later among the works it cites.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Depth with nonlinearity creates no bad local minima in resnets
Kenji Kawaguchi and Yoshua Bengio · 2019
Later among the works it cites.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
Kenji Kawaguchi and Jiaoyang Huang · 2019
Later among the works it cites.
Effect of depth and width on local minima in deep learning
Kenji Kawaguchi, Jiaoyang Huang, and Leslie Pack Kaelbling · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Adding one neuron can eliminate all bad local minima
Shiyu Liang, Ruoyu Sun, Jason D Lee, and R Srikant · 2018
Cited alongside, same era.
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Quynh Nguyen · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Graphmix: Regularized training of graph neural networks for semi-supervised learning
Vikas Verma, Meng Qu, Kenji Kawaguchi, Alex Lamb, Yoshua Bengio, Juho Kannala, and Jian Tang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Modeling from features: a mean-field framework for over-parameterized deep neural networks
Cong Fang, Jason D Lee, Pengkun Yang, and Tong Zhang · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
Elimination of all bad local minima in deep learning
Kenji Kawaguchi and Leslie Kaelbling · 2020
Later among the works it cites.
Andrea Montanari and Yiqiao Zhong · 2020
Later among the works it cites.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Suriya Gunasekar, Blake Woodworth, Jason D Lee, Nathan Srebro, and Daniel Soudry · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
A note on connectivity of sublevel sets in deep learning
Quynh Nguyen · 2021
Closest in time.