Fetching the paper…
Reading the bibliography…
We develop fast algorithms and robust software for convex optimization of two-layer neural networks with ReLU activation functions.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
MINOS 5.0 user’s guide
Murtagh, B. A. and Saunders, M. A · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) {O}(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Training a 3-node neural network is NP-Complete
Blum, A. and Rivest, R. L · 1988
Earlier work this paper cites.
On the convergence of the proximal point algorithm for convex minimization
Güler, O · 1991
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
Boser, B. E., Guyon, I., and Vapnik, V · 1992
Earlier work this paper cites.
A generalized theorem of the maximum
Ausubel, L. M. and Deneckere, R. J · 1993
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming , volume 13 of Siam studies in applied mathematics
Nesterov, Y. E. and Nemirovskii, A · 1994
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. J · 1999
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Dolan, E. D. and Moré, J. J · 2002
Earlier work this paper cites.
Linear semi-infinite programming theory: An updated survey
Goberna, M. A. and López, M. A · 2002
Earlier work this paper cites.
Convex neural networks
Bengio, Y., Roux, N. L., Vincent, P., Delalleau, O., and Marcotte, P · 2006
Earlier work this paper cites.
Extreme learning machine: Theory and applications
Huang, G., Zhu, Q., and Siew, C. K · 2006
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Beck, A. and Teboulle, M · 2009
Earlier work this paper cites.
Convex optimization theory , volume 1
Bertsekas, D · 2009
Earlier work this paper cites.
Efficient online and batch learning using forward backward splitting
Duchi, J. C. and Singer, Y · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Large-scale sparse logistic regression
Liu, J., Chen, J., and Ye, J · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J. C., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Convergence rates of inexact proximal-gradient methods for convex optimization
Schmidt, M., Le Roux, N., and Bach, F. R · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y · 2012
Cited alongside, same era.
Optimization for machine learning
Sra, S., Nowozin, S., and Wright, S. J · 2012
Cited alongside, same era.
LANCELOT: a Fortran package for large-scale nonlinear optimization (Release A) , volume 17
Conn, A. R., Gould, G., and Toint, P. L · 2013
Cited alongside, same era.
Gradient methods for minimizing composite functions
Nesterov, Y. E · 2013
Cited alongside, same era.
Constrained optimization and Lagrange multiplier methods
Bertsekas, D. P · 2014
Cited alongside, same era.
Practical augmented Lagrangian methods for constrained optimization , volume 10 of Fundamentals of algorithms
Birgin, E. G. and Martínez, J. M · 2014
Cited alongside, same era.
Geometry of optimization and implicit regularization in deep learning
Neyshabur, B., Tomioka, R., Salakhutdinov, R., and Srebro, N · 2017
Later among the works it cites.
A rewriting system for convex optimization problems
Agrawal, A., Verschueren, R., Diamond, S., and Boyd, S · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
MOSEK Optimizer API for Python 9.3.6 , 2019
ApS, M · 2019
Later among the works it cites.
Greedy layerwise learning can scale to ImageNet
Belilovsky, E., Eickenberg, M., and Oyallon, E · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do we need hundreds of classifiers to solve real world classification problems?
Delgado, M. F., Cernadas, E., Barro, S., and Amorim, D. G · 2014
Cited alongside, same era.
Monotonicity and restart in fast gradient methods
Giselsson, P. and Boyd, S. P · 2014
Cited alongside, same era.
Proximal algorithms
Parikh, N. and Boyd, S. P · 2014
Cited alongside, same era.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Cited alongside, same era.
Inexact accelerated augmented Lagrangian methods
Kang, M., Kang, M., and Jung, M · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Later among the works it cites.
Decoupling gating from linearity
Fiat, J., Malach, E., and Shalev-Shwartz, S · 2019
Later among the works it cites.
Choosing the step size: Intuitive line search algorithms with efficient convergence
Fridovich-Keil, S. and Recht, B · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
A brief prehistory of double descent
Loog, M., Viering, T., Mey, A., Krijthe, J. H., and Tax, D. M · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2020
Later among the works it cites.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks
Pilanci, M. and Ergen, T · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B. E., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
On the reproducibility of neural network predictions
Bhojanapalli, S., Wilber, K., Veit, A., Rawat, A. S., Kim, S., Menon, A. K., and Kumar, S · 2021
Later among the works it cites.
Global optimality beyond two layers: Training deep ReLU networks via convex programs
Ergen, T. and Pilanci, M · 2021
Later among the works it cites.
Implicit convex regularizers of CNN architectures: Convex optimization of two- and three-layer networks in polynomial time
Ergen, T. and Pilanci, M · 2021
Later among the works it cites.
Demystifying batch normalization in relu networks: Equivalent convex optimization models and implicit regularization
Ergen, T., Sahiner, A., Ozturkler, B., Pauly, J. M., Mardani, M., and Pilanci, M · 2021
Later among the works it cites.
Exact and relaxed convex formulations for shallow neural autoregressive models
Gupta, V., Bartan, B., Ergen, T., and Pilanci, M · 2021
Later among the works it cites.
Vector-output ReLU neural network problems are copositive programs: Convex analysis of two layer networks and polynomial-time algorithms
Sahiner, A., Ergen, T., Pauly, J. M., and Pilanci, M · 2021
Later among the works it cites.
The curious case of convex neural networks
Sivaprasad, S., Singh, A., Manwani, N., and Gandhi, V · 2021
Later among the works it cites.
The hidden convex optimization landscape of regularized two-layer relu networks: an exact characterization of optimal solutions
Wang, Y., Lacotte, J., and Pilanci, M · 2021
Later among the works it cites.
Bai, Y., Gautam, T., and Sojoudi, S · 2022
Closest in time.