Fetching the paper…
Reading the bibliography…
We study regularized deep neural networks (DNNs) and introduce a convex analytic framework to characterize the structure of the hidden layers.
How do infinite width bounded norm networks look in function space?
Savarese, P., Evron, I., Soudry, D., and Srebro, N · 1902
Earlier work this paper cites.
Linear semi-infinite optimization
Goberna, M. A. and López-Cerdá, M · 1998
Earlier work this paper cites.
Convex geometry and duality of over-parameterized neural networks
Ergen, T. and Pilanci, M · 2002
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
L1 regularization in infinite dimensional feature spaces
Rosset, S., Swirszcz, G., Srebro, N., and Zhu, J · 2007
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
The CIFAR-10 dataset
Krizhevsky, A., Nair, V., and Hinton, G · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Cited alongside, same era.
A closer look at few-shot classification
Chen, W.-Y., Liu, Y.-C., Kira, Z., Wang, Y.-C. F., and Huang, J.-B · 2018
Cited alongside, same era.
Decorrelated batch normalization
Huang, L., Yang, D., Lang, B., and Deng, J · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Laurent, T. and Brecht, J · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M · 2019
Later among the works it cites.
Minimum “norm” neural networks are splines
Parhi, R. and Nowak, R. D · 2019
Later among the works it cites.
Deep neural networks with multi-branch architectures are intrinsically less non-convex
Zhang, H., Shao, J., and Salakhutdinov, R · 2019
Later among the works it cites.
Framelet pooling aided deep learning network: the method to process high dimensional medical data
Hyun, C. M., Kim, K. C., Cho, H. C., Choi, J. K., and Seo, J. K · 2020
Closest in time.
Prevalence of neural collapse during the terminal phase of deep learning training
Papyan, V., Han, X., and Donoho, D. L · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shamir, O · 2018
Cited alongside, same era.
On the margin theory of feedforward neural networks
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Du, S. and Hu, W · 2019
Cited alongside, same era.
Convex duality and cutting plane methods for over-parameterized neural networks
Ergen, T. and Pilanci, M · 2019
Cited alongside, same era.
Convex optimization for shallow neural networks
Ergen, T. and Pilanci, M · 2019
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Arora, S., Cohen, N., Golowich, N., and Hu, W
Cited in the paper.
Pilanci, M. and Ergen, T · 2020
Closest in time.
Fsnet: Feature selection network on high-dimensional biological data
Singh, D. and Yamada, M · 2020
Closest in time.
Implicit convex regularizers of cnn architectures: Convex optimization of two- and three-layer networks in polynomial time
Ergen, T. and Pilanci, M · 2021
Closest in time.
Ergen, T., Sahiner, A., Ozturkler, B., Pauly, J. M., Mardani, M., and Pilanci, M · 2021
Closest in time.
Convex neural autoregressive models: Towards tractable, expressive, and theoretically-backed models for sequential forecasting and generation
Gupta, V., Bartan, B., Ergen, T., and Pilanci, M · 2021
Closest in time.
Vector-output relu neural network problems are copositive programs: Convex analysis of two layer networks and polynomial-time algorithms
Sahiner, A., Ergen, T., Pauly, J. M., and Pilanci, M · 2021
Closest in time.