Fetching the paper…
Reading the bibliography…
Understanding the global optimality in deep learning (DL) has been attracting more and more attention recently.
Training a 3-node neural network is np-complete
A. Blum and R. L. Rivest · 1989
Earlier work this paper cites.
The mnist database of handwritten digits
Y. LeCun · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A branch-and-prune method for global optimization
D. G. Sotiropoulos and T. N. Grapsa · 2001
Earlier work this paper cites.
Applied Mathematics Body and Soul: Vol I-III
K. Eriksson, D. Estep, and C. Johnson · 2003
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2011 (VOC2011) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Later among the works it cites.
Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
I. Panageas and G. Piliouras · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matconvnet: Convolutional neural networks for matlab
A. Vedaldi and K. Lenc · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, and Y. LeCun · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Closest in time.
Global optimality in neural network training
B. D. Haeffele and R. Vidal · 2017
Closest in time.
Global guarantees for enforcing deep generative priors by empirical risk
P. Hand and V. Voroninski · 2017
Closest in time.
Global optimization of lipschitz functions
C. Malherbe and N. Vayatis · 2017
Closest in time.
Variants of rmsprop and adagrad with logarithmic regret bounds
M. C. Mukkamala and M. Hein · 2017
Closest in time.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Closest in time.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2017
Closest in time.
Convergent block coordinate descent for training tikhonov regularized deep neural networks
Z. Zhang and M. Brand · 2017
Closest in time.