Fetching the paper…
Reading the bibliography…
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Understanding dropout
Baldi, P. and Sadowski, P. J · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P. S · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Deeppose: Human pose estimation via deep neural networks
Toshev, A. and Szegedy, C · 2014
Earlier work this paper cites.
Follow the leader with dropout perturbations
Van Erven, T., Kotłowski, W., and Warmuth, M. K · 2014
Earlier work this paper cites.
Altitude training: Strong bounds for single-layer dropout
Wager, S., Fithian, W., Wang, S., and Liang, P. S · 2014
Earlier work this paper cites.
On the inductive bias of dropout
Helmbold, D. P. and Long, P. M · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Dropout training of matrix factorization and autoencoder for link prediction in sparse graphs
Zhai, S. and Zhang, Z · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Dropout rademacher complexity of deep neural networks
Gao, W. and Zhou, Z.-H · 2016
Cited alongside, same era.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
Identity matters in deep learning
Hardt, M. and Ma, T · 2016
On the relationship between dropout and equiangular tight frames
Bank, D. and Giryes, R · 2018
Later among the works it cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Bartlett, P., Helmbold, D., and Long, P · 2018
Later among the works it cites.
Dropout as a low-rank regularizer for matrix factorization
Cavazza, J., Haeffele, B. D., Lane, C., Morerio, P., Murino, V., and Vidal, R · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M · 2018
Later among the works it cites.
Deep linear networks with arbitrary loss: All local minima are global
Laurent, T. and Brecht, J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dropout non-negative matrix factorization for independent feature learning
He, Z., Liu, J., Liu, C., Wang, Y., Yin, A., and Huang, Y · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Cited alongside, same era.
Surprising properties of dropout in deep networks
Helmbold, D. P. and Long, P. M · 2017
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Arora, S., Cohen, N., Golowich, N., and Hu, W · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Later among the works it cites.
Martin, C. H. and Mahoney, M. W · 2018
Later among the works it cites.
On the implicit bias of dropout
Mianjy, P., Arora, R., and Vidal, R · 2018
Later among the works it cites.
Dropout training, data-dependent regularization, and generalization bounds
Mou, W., Zhou, Y., Gao, J., and Wang, L · 2018
Later among the works it cites.
Convergence of gradient descent on separable data
Nacson, M. S., Lee, J., Gunasekar, S., Srebro, N., and Soudry, D · 2018
Later among the works it cites.
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
Azizan, N. and Hassibi, B · 2019
Closest in time.
A comprehensive analysis of deep regression
Lathuilière, S., Mesejo, P., Alameda-Pineda, X., and Horaud, R · 2019
Closest in time.