Fetching the paper…
Reading the bibliography…
We present a novel layerwise optimization algorithm for the learning objective of Piecewise-Linear Convolutional Neural Networks (PL-CNNs), a large class of convolutional neural networks.
On the expressibility of piecewise-linear continuous functions as the difference of two piecewise-linear convex functions
D. Melzer · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
David Rumelhart, Geoffrey Hinton, and Ronald Williams · 1986
Earlier work this paper cites.
DC programming: overview
Reiner Horst and Nguyen V. Thoai · 1999
Earlier work this paper cites.
The concave-convex procedure (CCCP)
Alan L. Yuille and Anand Rangarajan · 2002
Earlier work this paper cites.
Support vector machine learning for interdependent and structured output spaces
Ioannis Tsochantaridis, Thomas Hofmann, Thorsten Joachims, and Yasemin Altun · 2004
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, Hugo Larochelle, et al · 2007
Earlier work this paper cites.
Cutting-plane training of structural SVMs
Thorsten Joachims, Thomas Finley, and Chun-Nam John Yu · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for SVM
Shai Shalev-Shwartz, Yoram Singer, and Nathan Srebro · 2009
Earlier work this paper cites.
Learning structural SVMs with latent variables
Chun-Nam John Yu and Thorsten Joachims · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Theano: new features and speed improvements, 2012
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian J. Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio · 2012
Cited alongside, same era.
Training deep and recurrent networks with hessian-free optimization
James Martens and Ilya Sutskever · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
Matthew Zeiler · 2012
Cited alongside, same era.
Block-coordinate Frank-Wolfe optimization for structural SVMs
Simon Lacoste-Julien, Martin Jaggi, Mark Schmidt, and Patrick Pletscher · 2013
Cited alongside, same era.
Riemannian metrics for neural networks
Yann Ollivier · 2013
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Later among the works it cites.
A multi-plane block-coordinate Frank-Wolfe algorithm for training structural SVMs with a costly max-oracle
Neel Shah, Vladimir Kolmogorov, and Christoph H. Lampert · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Later among the works it cites.
Brandon Amos, Lei Xu, and J. Zico Kolter · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Proximal algorithms
Neal Parikh and Stephen Boyd · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Difference target propagation
Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio · 2015
Cited alongside, same era.
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2016
Closest in time.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Closest in time.
Improper deep kernels
Uri Heinemann, Roi Livni, Elad Eban, Gal Elidan, and Amir Globerson · 2016
Closest in time.
Partial linearization based optimization for multi-class SVM
Pritish Mohapatra, Puneet Dokania, CV Jawahar, and M Pawan Kumar · 2016
Closest in time.
Minding the gaps for block Frank-Wolfe optimization of structured SVMs
Anton Osokin, Jean-Baptiste Alayrac, Isabella Lukasewitz, Puneet Dokania, and Simon Lacoste-Julien · 2016
Closest in time.
Training neural networks without gradients: A scalable ADMM approach
Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein · 2016
Closest in time.
Convexified convolutional neural networks
Yuchen Zhang, Percy Liang, and Martin J. Wainwright · 2016
Closest in time.