Fetching the paper…
Reading the bibliography…
We develop a convex analytic approach to analyze finite width two-layer ReLU networks.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 1904
Earlier work this paper cites.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
Principles of Mathematical Analysis
Walter Rudin · 1964
Earlier work this paper cites.
Convex Analysis
R. T. Rockafellar · 1970
Earlier work this paper cites.
On milman’s inequality and random subspaces which escape through a mesh in rn
Yehoram Gordon · 1988
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Mcgraw-hill science
Tom M Mitchell and Machine Learning · 1997
Earlier work this paper cites.
Linear semi-infinite optimization
Miguel Angel Goberna and Marco López-Cerdá · 1998
Earlier work this paper cites.
Sdpt3—a matlab software package for semidefinite-quadratic-linear programming, version 3.0
RH Tütüncü, KC Toh, and MJ Todd · 2001
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Decoding by linear programming
E. J. Candes and T. Tao · 2005
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2006
Earlier work this paper cites.
For most large underdetermined systems of linear equations, the minimal ℓ 1 \ell_{1} -norm solution is also the sparsest solution
D. L. Donoho · 2006
Earlier work this paper cites.
On model selection consistency of lasso
Peng Zhao and Bin Yu · 2006
Earlier work this paper cites.
Particular formulae for the moore–penrose inverse of a columnwise partitioned matrix
Jerzy K Baksalary and Oskar Maria Baksalary · 2007
Earlier work this paper cites.
L1 regularization in infinite dimensional feature spaces
Saharon Rosset, Grzegorz Swirszcz, Nathan Srebro, and Ji Zhu · 2007
Earlier work this paper cites.
SPGL1: A solver for large-scale sparse reconstruction, June 2007
E. van den Berg and M. P. Friedlander · 2007
Earlier work this paper cites.
Convex and semi-nonnegative matrix factorizations
Chris HQ Ding, Tao Li, and Michael I Jordan · 2008
Cited alongside, same era.
Hardness of learning halfspaces with noise
Venkatesan Guruswami and Prasad Raghavendra · 2009
Cited alongside, same era.
Equivalence of minimal ℓ 0 \ell_{0} -and ℓ p \ell_{p} -norm solutions of linear equalities, inequalities and linear programs for sufficiently small p
GM Fung and OL Mangasarian · 2011
Cited alongside, same era.
Learning feature representations with k-means
Adam Coates and Andrew Y Ng · 2012
Cited alongside, same era.
Probability in Banach Spaces: isoperimetry and processes
Michel Ledoux and Michel Talagrand · 2013
Cited alongside, same era.
CVX: Matlab software for disciplined convex programming, version 2.1
Michael Grant and Stephen Boyd · 2014
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Later among the works it cites.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Convex relaxations of convolutional neural nets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The CIFAR-10 dataset
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Burak Bartan and Mert Pilanci · 2019
Later among the works it cites.
On the inductive bias of neural tangent kernels
Alberto Bietti and Julien Mairal · 2019
Later among the works it cites.
Convex optimization for shallow neural networks
Tolga Ergen and Mert Pilanci · 2019
Later among the works it cites.
Minimum “norm” neural networks are splines
Rahul Parhi and Robert D Nowak · 2019
Later among the works it cites.
Supermann: a superlinearly convergent algorithm for finding fixed points of nonexpansive operators
Andreas Themelis and Panagiotis Patrinos · 2019
Later among the works it cites.
Gradient dynamics of shallow univariate relu networks
Francis Williams, Matthew Trager, Claudio Silva, Daniele Panozzo, Denis Zorin, and Joan Bruna · 2019
Later among the works it cites.
Jonathan Lacotte and Mert Pilanci · 2020
Closest in time.
No spurious local minima: on the optimization landscapes of wide and deep neural networks, 2020
Johannes Lederer · 2020
Closest in time.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks
Mert Pilanci and Tolga Ergen · 2020
Closest in time.
Neural kernels without tangents
Vaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil, Jonathan Ragan-Kelley, Ludwig Schmidt, and Benjamin Recht · 2020
Closest in time.
Implicit convex regularizers of {cnn} architectures: Convex optimization of two- and three-layer networks in polynomial time
Tolga Ergen and Mert Pilanci · 2021
Closest in time.
Tolga Ergen, Arda Sahiner, Batu Ozturkler, John M. Pauly, Morteza Mardani, and Mert Pilanci · 2021
Closest in time.
Convex neural autoregressive models: Towards tractable, expressive, and theoretically-backed models for sequential forecasting and generation
Vikul Gupta, Burak Bartan, Tolga Ergen, and Mert Pilanci · 2021
Closest in time.
The unreasonable effectiveness of patches in deep convolutional kernels methods
Louis Thiry, Michael Arbel, Eugene Belilovsky, and Edouard Oyallon · 2021
Closest in time.