Fetching the paper…
Reading the bibliography…
We formalize the notion of a pseudo-ensemble, a (possibly infinite) collection of child models spawned from a parent model by perturbing it according to some noise process.
Semi-Supervised Learning
Y. Grandvalet and Y. Bengio · 2006
Earlier work this paper cites.
Elements of Statistical Learning II
T. Hastie, J. Friedman, and R. Tibshirani · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, and Y. Bengio · 2008
Earlier work this paper cites.
Deep learning via semi-supervised embedding
J. Weston, F. Ratle, and R. Collobert · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Lectures on Stochastic Programming: Modeling and Theory
A. Shapiro, D. Dentcheva, and A. Ruszczynski · 2009
Earlier work this paper cites.
Robust regression and lasso
H. Xu, C. Caramanis, and S. Mannor · 2009
Earlier work this paper cites.
Robustness and regularization of support vector machines
H. Xu, C. Caramanis, and S. Mannor · 2009
Earlier work this paper cites.
Theano: A cpu and gpu math expression compiler
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio · 2010
Earlier work this paper cites.
Theory and applications of robust optimization
D. Bertsimas, D. B. Brown, and C. Caramanis · 2011
Cited alongside, same era.
Workshop on challenges in learning hierarchical models: Transfer learning and optimization
Q. V. Le, M. A. Ranzato, R. R. Salakhutdinov, A. Y. Ng, and J. Tenenbaum · 2011
Cited alongside, same era.
The manifold tangent classifier
S. Rifai, Y. Dauphin, P. Vincent, Y. Bengio, and X. Muller · 2011
Cited alongside, same era.
Large-scale feature learning with spike-and-slab sparse coding
I. J. Goodfellow, A. Courville, and Y. Bengio · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
G.E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R.R. Salakhutdinov · 2012
Cited alongside, same era.
Understanding dropout
P. Baldi and P. Sadowski · 2013
On the difficulties of training recurrent neural networks
R. Pacanu, T. Mikolov, and Y. Bengio · 2013
Later among the works it cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Y. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Later among the works it cites.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. Liang · 2013
Later among the works it cites.
Deep generative stochastic networks trainable by backprop
Y. Bengio, É. Thibodeau-Laufer, G. Alain, and J. Yosinski · 2014
Closest in time.
A convolutional neural network for modelling sentences
N. Kalchbrenner, E. Grefenstette, and P. Blunsom · 2014
Closest in time.
Distributed representations of sentences and documents
Q. V. Le and T. Mikolov · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning with marginalized corrupted features
L. Van der Maaten, M. Chen, S. Tyree, and K. Q. Weinberger · 2013
Cited alongside, same era.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
D.-H. Lee · 2013
Cited alongside, same era.
Closest in time.
Learning ordered representations with nested dropout
O. Rippel, M. A. Gelbart, and R. P. Adams · 2014
Closest in time.
An empirical analysis of dropout in piecewise linear networks
D. Warde-Farley, I. J. Goodfellow, A. Courville, and Y. Bengio · 2014
Closest in time.