Improved sample complexities for deep networks and robust classification via an all-layer margin
Original
Wei, C. and Ma, T · 1910
Earlier work this paper cites.
The penn treebank: annotating predicate argument structure
Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B · 1994
Earlier work this paper cites.
Concentration inequalities and empirical processes theory applied to the analysis of learning algorithms
Bousquet, O · 2002
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Zhang, T · 2002
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Kakade, S. M., Sridharan, K., and Tewari, A · 2009
Earlier work this paper cites.
Self-concordant analysis for logistic regression
Bach, F. et al · 2010
Earlier work this paper cites.
Smoothness, low noise and fast rates
Srebro, N., Sridharan, K., and Tewari, A · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Original
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Understanding dropout
Baldi, P. and Sadowski, P. J · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P. S · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Fast dropout training
Wang, S. and Manning, C · 2013
Earlier work this paper cites.
A bayesian encourages dropout
Original
Maeda, S.-i · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Altitude training: Strong bounds for single-layer dropout
Wager, S., Fithian, W., Wang, S., and Liang, P. S · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Original
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
On the inductive bias of dropout
Helmbold, D. P. and Long, P. M · 2015
Earlier work this paper cites.
Regularizing rnns by stabilizing activations
Original
Krueger, D. and Memisevic, R · 2015
Earlier work this paper cites.
Quasi-recurrent neural networks
Original
Bradbury, J., Merity, S., Xiong, C., and Socher, R · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Original
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Dropout with expectation-linear regularization
Original
Ma, X., Gao, Y., Hu, Z., Yu, Y., Deng, Y., and Hovy, E · 2016
Earlier work this paper cites.