Fetching the paper…
Reading the bibliography…
Dropout and other feature noising schemes control overfitting by artificially corrupting the training data.
Learning from hints in neural networks
Yaser S Abu-Mostafa · 1990
Earlier work this paper cites.
Noise injection into inputs in back-propagation learning
Kiyotoshi Matsuoka · 1992
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
Chris M Bishop · 1995
Earlier work this paper cites.
Improving the accuracy and speed of support vector machines
Chris J.C. Burges and Bernhard Schölkopf · 1997
Earlier work this paper cites.
Theory of Point Estimation
Erich Leo Lehmann and George Casella · 1998
Earlier work this paper cites.
Transductive inference for text classification using support vector machines
Thorsten Joachims · 1999
Earlier work this paper cites.
Transformation invariance in pattern recognition: Tangent distance and propagation
Patrice Y Simard, Yann A Le Cun, John S Denker, and Bernard Victorri · 2000
Earlier work this paper cites.
Text classification from labeled and unlabeled documents using EM
Kamal Nigam, Andrew Kachites McCallum, Sebastian Thrun, and Tom Mitchell · 2000
Earlier work this paper cites.
The trade-off between generative and discriminative classifiers
G. Bouchard and B. Triggs · 2004
Earlier work this paper cites.
Classification with hybrid generative/discriminative models
R. Raina, Y. Shen, A. Ng, and A. McCallum · 2004
Cited alongside, same era.
Entropy regularization
Y. Grandvalet and Y. Bengio · 2005
Cited alongside, same era.
Semi-supervised structured output learning based on a hybrid generative and discriminative approach
J. Suzuki, A. Fujino, and H. Isozaki · 2007
Cited alongside, same era.
Adaptive regularization of weight vectors
Koby Crammer, Alex Kulesza, Mark Dredze, et al · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2010
Cited alongside, same era.
Regularization paths for generalized linear models via coordinate descent
Jerome Friedman, Trevor Hastie, and Rob Tibshirani · 2010
Cited alongside, same era.
Learning word vectors for sentiment analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Later among the works it cites.
Large scale text classification using semi-supervised multinomial naive Bayes
Jiang Su, Jelber Sayyad Shirab, and Stan Matwin · 2011
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Later among the works it cites.
Baselines and bigrams: Simple, good sentiment and topic classification
Sida Wang and Christopher D Manning · 2012
Later among the works it cites.
Learning with marginalized corrupted features
Laurens van der Maaten, Minmin Chen, Stephen Tyree, and Kilian Q Weinberger · 2013
Closest in time.
Fast dropout training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The manifold tangent classifier
Salah Rifai, Yann Dauphin, Pascal Vincent, Yoshua Bengio, and Xavier Muller · 2011
Cited alongside, same era.
Adding noise to the input of a model trained with a regularized objective
Salah Rifai, Xavier Glorot, Yoshua Bengio, and Pascal Vincent · 2011
Cited alongside, same era.
Sida I Wang and Christopher D Manning · 2013
Closest in time.
Feature noising for log-linear structured prediction
Sida I Wang, Mengqiu Wang, Stefan Wager, Percy Liang, and Christopher D Manning · 2013
Closest in time.
Maxout networks
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Closest in time.