Fetching the paper…
Reading the bibliography…
Several recently proposed architectures of neural networks such as ResNeXt, Inception, Xception, SqueezeNet and Wide ResNet are based on the designing idea of having multiple branches and have demonstrated improved performance in many applications.
Quasi-equilibria in markets with non-convex preferences
R. M. Starr · 1969
Earlier work this paper cites.
Generalized linear programming solves the dual
T. L. Magnanti, J. F. Shapiro, and M. H. Wagner · 1976
Earlier work this paper cites.
Estimates of the duality gap for large-scale separable nonconvex optimization problems
D. P. Bertsekas and N. R. Sandell · 1982
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
A. Blum and R. L. Rivest · 1989
Earlier work this paper cites.
On the sum of the largest eigenvalues of a symmetric matrix
M. L. Overton and R. S. Womersley · 1992
Earlier work this paper cites.
On the complexity of training neural networks with continuous activation functions
B. DasGupta, H. T. Siegelmann, and E. Sontag · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Hardness results for neural network approximation problems
P. Bartlett and S. Ben-David · 1999
Earlier work this paper cites.
Strong duality in nonconvex quadratic optimization with two quadratic constraints
A. Beck and Y. C. Eldar · 2006
Earlier work this paper cites.
Convex neural networks
Y. Bengio, N. L. Roux, P. Vincent, O. Delalleau, and P. Marcotte · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Complex-valued autoencoders
P. Baldi and Z. Lu · 2012
Earlier work this paper cites.
On the O ( 1 / n ) O(1/n) convergence rate of the Douglas–Rachford alternating direction method
B. He and X. Yuan · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2013
Earlier work this paper cites.
Convex deep learning via normalized kernels
Ö. Aslan, X. Zhang, and D. Schuurmans · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Open problem: The landscape of the loss surfaces of multilayer networks
A. Choromanska, Y. LeCun, and G. B. Arous · 2015
Earlier work this paper cites.
Blessing of massive scale: Spatial graphical model estimation with a total cardinality constraint
E. X. Fang, H. Liu, and M. Wang · 2015
Earlier work this paper cites.
Escaping from saddle points-online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
B. Neyshabur, R. Salakhutdinov, and N. Srebro · 2015
Earlier work this paper cites.
Deep linear neural networks: A theory of learning in the brain and mind
A. M. Saxe · 2015
Earlier work this paper cites.
Efficient approaches for escaping higher order saddle points in non-convex optimization
A. Anandkumar and R. Ge · 2016
Cited alongside, same era.
Learning and 1-bit compressed sensing under asymmetric noise
P. Awasthi, M.-F. Balcan, N. Haghtalab, and H. Zhang · 2016
Cited alongside, same era.
Refined Shapely-Folkman lemma and its application in duality gap estimation
Y. Bi and A. Tang · 2016
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
Low-rank optimization with convex constraints
C. Grussler, A. Rantzer, and P. Giselsson · 2016
Cited alongside, same era.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 MB model size
The multilinear structure of ReLU networks
T. Laurent and J. von Brecht · 2017
Later among the works it cites.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, and T. Goldstein · 2017
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2017
Later among the works it cites.
Depth creates no bad local minima
H. Lu and K. Kawaguchi · 2017
Later among the works it cites.
Spurious local minima are common in two-layer ReLU neural networks
I. Safran and O. Shamir · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Bounding duality gap for separable problems with linear constraints
M. Udell and S. Boyd · 2016
Cited alongside, same era.
Generalized low rank models
M. Udell, C. Horn, R. Zadeh, and S. Boyd · 2016
Cited alongside, same era.
Residual networks behave like ensembles of relatively shallow networks
A. Veit, M. J. Wilber, and S. Belongie · 2016
Cited alongside, same era.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
Failures of gradient-based deep learning
S. Shalev-Shwartz, O. Shamir, and S. Shammah · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2017
Later among the works it cites.
Inception-v4, Inception-ResNet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2017
Later among the works it cites.
R. Vidal, J. Bruna, R. Giryes, and S. Soatto · 2017
Later among the works it cites.
Diverse neural network learns true target functions
B. Xie, Y. Liang, and L. Song · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
On the learnability of fully-connected neural networks
Y. Zhang, J. Lee, M. Wainwright, and M. Jordan · 2017
Later among the works it cites.
Convexified convolutional neural networks
Y. Zhang, P. Liang, and M. J. Wainwright · 2017
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Closest in time.
Matrix completion and related problems via strong duality
M.-F. Balcan, Y. Liang, D. P. Woodruff, and H. Zhang · 2018
Closest in time.
SGD learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Closest in time.
On the power of over-parametrization in neural networks with quadratic activation
S. S. Du and J. D. Lee · 2018
Closest in time.
Deep semi-random features for nonlinear function approximation
K. Kawaguchi, B. Xie, and L. Song · 2018
Closest in time.
Adding one neuron can eliminate all bad local minima
S. Liang, R. Sun, J. D. Lee, and R. Srikant · 2018
Closest in time.
Understanding the loss surface of neural networks for binary classification
S. Liang, R. Sun, Y. Li, and R. Srikant · 2018
Closest in time.
Towards understanding the role of over-parametrization in generalization of neural networks
B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro · 2018
Closest in time.
Empirical risk landscape analysis for understanding deep neural networks
P. Zhou and J. Feng · 2018
Closest in time.