Fetching the paper…
Reading the bibliography…
Noise injection (NI) is an efficient technique to mitigate over-fitting in neural networks (NNs).
A. Ivakhnenko, “Polynomial theory of complex systems,” IEEE Transactions on Systems, Man and Cybernetics , vol. 4, pp. 364–378, 1971
1971
Earlier work this paper cites.
D. C. Plaut, S. J. Nowlan, and G. E. Hinton, “Experiments on learning by back-propagation,” Tech. Rep. CMU-CS-86-126 , 1986
1986
Earlier work this paper cites.
J. Sietsma and R. J. F. Dow, “Neural network pruning - why and how,” Proceedings of the IEEE International Conferences on Neural Networks , vol. 1, pp. 325–333, 1988
1988
Earlier work this paper cites.
M. C. Mozer and P. Smolensky, “Skeletonization: A technique for trimming the fat from a network via relevance assessment,” 1989, pp. 107–115
1989
Earlier work this paper cites.
S. J. Hanson and L. Y. Pratt, “Comparing biases for minimal network construction with back-propagation,” 1989, pp. 177–185
1989
Earlier work this paper cites.
Y. L. Cun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” 1990, pp. 598–605
1990
Earlier work this paper cites.
K. Hornik, “Approximation capabilities of multilayer feedforward networkss,” Neural Networks, , vol. 4, no. 2, pp. 251–257, 1991
1991
Earlier work this paper cites.
L. Holmstrom and P. Koistinen, “Using additive noise in back-propagation training,” IEEE transactions on neural networks , vol. 3, no. 1, pp. 24–38, 1992
1992
Earlier work this paper cites.
K. Matsuoka, “Noise injection into inputs in back-propagation learning,” IEEE Tranactions on Systems, Man, and Cybernetics , vol. 22, no. 3, pp. 436–440, 1992
1992
Earlier work this paper cites.
I. Frank and J. Friedman, “A statistical view of some chemometrics regression tools,” Technometrics , vol. 35, no. 2, pp. 109–135, 1993
1993
Earlier work this paper cites.
C. M. Bishop, “Training with noise is equivalent to Tikhonov regularization,” Neural Computation , vol. 7, no. 1, pp. 108–116, 1995
1995
Earlier work this paper cites.
G. Lugosi, “Nonparametric estimation via empirical risk minimization,” IEEE Transactions on Information Theory , vol. 41, 1995
1995
Earlier work this paper cites.
G. An, “The Effects of Adding Noise During Backpropagation Training on a Generalization Performance,” Neural Computation , vol. 8, pp. 643–674, 1996
1996
Earlier work this paper cites.
Y. Grandvalet, S. Canu, and S. Boucheron, “Noise Injection: Theoretical Prospects,” Neural Computation , vol. 9, no. 5, pp. 1093–1108, 1997
1997
Cited alongside, same era.
N. Srebro and A. Shraibman, “Rank, trace-norm and max-norm,” Proceedings of the 18th annual conference on Learning Theory, COLT’05 , pp. 545–560, 2005
2005
Cited alongside, same era.
T. Zhang and B. Yu, “Boosting with early stopping: Convergence and consistency,” The Annals of Statistics , vol. 33, no. 4, pp. 1538–1579, 2005
2005
Cited alongside, same era.
H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,” Journal of the Royal Statistical Society: Series B , vol. 67, no. 2, pp. 301–320, 2005
2005
Cited alongside, same era.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science , vol. 313, pp. 504–507, 2006
2012
Later among the works it cites.
L. Wan, M. Zeiler, S. Zhang, Y. LeCun, and R. Fergus, “Regularization of neural networks using dropConnect,” Proceedings of Machine Learning Research , vol. 28, no. 3, pp. 1058–1066, 2013
2013
Later among the works it cites.
I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, “Maxout Networks,” Proceedings of the 30th International Conference on Machine Learning (ICML) , vol. 28, pp. 1319–1327, 2013
2013
Later among the works it cites.
S. Wang and C. D. Manning, “Fast dropout training,” Proceedings of the 30th International Conference on Machine Learning , vol. 28, pp. 118–126, 2013
2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2006
Cited alongside, same era.
R. Tibshirani, “Regression shrinkage and selection via the lasso,” Journal of the Royal Statistical Society: Series B , vol. 58, no. 1, pp. 267–288, 2006
2006
Cited alongside, same era.
H. Zou, “The adaptive lasso and its oracle properties,” Journal of the American Statistical Association: Theory and Methods , vol. 101, no. 476, pp. 1418–1429, 2006
2006
Cited alongside, same era.
M. Yuan and Y. Lin, “Model selection and estimation in regression with grouped variables,” Journal of the Royal Statistical Society: Series B , vol. 68, no. 1, pp. 49–67, 2006
2006
Cited alongside, same era.
Y. Yao, L. Rosasco, and A. Caponnetto, “On early stopping in gradient descent learning,” Constructive Approximation , vol. 26, no. 2, pp. 289–315, 2007
2007
Cited alongside, same era.
M. J. Wainwright and M. I. Jordan, Graphical Models, Exponential Families, and Variational Inference . Now Publishers Inc, 2008
2008
Cited alongside, same era.
2008
Cited alongside, same era.
——, “Hand movement recognition for brazilian sign language: A study using distance-based neural networks,” Proceedings of the 2009 International Joint Conference on Neural Networks , pp. 2355–2362, 2009
2009
Cited alongside, same era.
J. Ba and B. Frey, “Adaptive dropout for training deep neural networks,” Advances in Neural Information Processing Systems , pp. 1–9, 2013
2013
Later among the works it cites.
S. Wager, S. Wang, and P. Liang, “Dropout training as adaptive regularization,” Advances in Neural Information Processing Systems (NIPS) , vol. 26, pp. 351–359, 2013
2013
Later among the works it cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, p. 1929−1958, 2014
2014
Later among the works it cites.
A. Tsanas, M. A. Little, C. Fox, and L. O. Ramig, “Objective automatic assessment of rehabilitative speech treatment in parkinson’s disease,” IEEE Transaction in Neural Systems and Rehabilitation Engineering , vol. 22, no. 1, pp. 181–190, 2014
2014
Later among the works it cites.
G. Kang, J. Li, and D. Tao, “Shakeout: A new regularized deep neural network training scheme,” Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence , pp. 1751–1757, 2016
2016
Closest in time.
B. Wang and D. Klabjan, “Regularization for unsupervised deep neural nets,” Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence , 2017
2017
Closest in time.
D. B. Dias, R. C. B. Madeo, T. Rocha, H. H. Biscaro, and S. M. Peres, “LIBRAS Movement Database,” https://archive.ics.uci.edu/ml/machine-learning-databases/libras/movement_libras.names , 2009, [Online; accessed 22-July-2017]
2017
Closest in time.
M. J. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint . to appear, 2018
2018
Closest in time.