Fetching the paper…
Reading the bibliography…
We describe and analyze a new boosting algorithm for deep learning called SelfieBoost.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R.E. Schapire · 1995
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Y. LeCun and Y. Bengio · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Boosting neural networks
Holger Schwenk and Yoshua Bengio · 2000
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Earlier work this paper cites.
Scaling learning algorithms towards ai
Y. Bengio and Y. LeCun · 2007
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
M.A. Ranzato, F.J. Huang, Y.L. Boureau, and Y. Lecun · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Olivier Bousquet and Léon Bottou · 2008
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Y. Bengio · 2009
Cited alongside, same era.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
H. Lee, R. Grosse, R. Ranganath, and A.Y. Ng · 2009
Cited alongside, same era.
Deep learning via hessian-free optimization
James Martens · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Improving deep neural networks for lvcsr using rectified linear units and dropout
G. Dahl, T. Sainath, and G. Hinton · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
Training recurrent neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Building high-level features using large scale unsupervised learning
Q. V. Le, M.-A. Ranzato, R. Monga, M. Devin, G. Corrado, K. Chen, J. Dean, and A. Y. Ng · 2012
Cited alongside, same era.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Cited alongside, same era.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Cited alongside, same era.
Ilya Sutskever · 2013
Later among the works it cites.
Visualizing and understanding convolutional neural networks
M. Zeiler and R. Fergus · 2013
Later among the works it cites.
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Closest in time.
Minimizing the maximal loss: How and why
Shai Shalev-Shwartz and Yonatan Wexler · 2016
Closest in time.