Fetching the paper…
Reading the bibliography…
Using weight decay to penalize the L2 norms of weights in neural networks has been a standard training practice to regularize the complexity of networks.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A · 1992
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Bartlett, P. L · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Cited alongside, same era.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Cisse, M., Bojanowski, P., Grave, E., Dauphin, Y., and Usunier, N · 2017
Cited alongside, same era.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., and Srebro, N · 2018
Later among the works it cites.
Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients
Ross, A. S. and Doshi-Velez, F · 2018
Later among the works it cites.
Adversarially robust generalization requires more data
Schmidt, L., Santurkar, S., Tsipras, D., Talwar, K., and Madry, A · 2018
Later among the works it cites.
Adaptive residual networks for high-quality image restoration
Zhang, Y., Sun, L., Yan, C., Ji, X., and Dai, Q · 2018
Later among the works it cites.
Efficient and accurate estimation of lipschitz constants for deep neural networks
Fazlyab, M., Robey, A., Hassani, H., Morari, M., and Pappas, G · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Spectral norm regularization for improving the generalizability of deep learning
Yoshida, Y. and Miyato, T · 2017
Cited alongside, same era.
Gilmer, J., Metz, L., Faghri, F., Schoenholz, S. S., Raghu, M., Wattenberg, M., and Goodfellow, I · 2018
Cited alongside, same era.
Sparse dnns with improved adversarial robustness
Guo, Y., Zhang, C., Zhang, C., and Chen, Y · 2018
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Cited alongside, same era.
On the implicit bias of dropout
Mianjy, P., Arora, R., and Vidal, R · 2018
Cited alongside, same era.
Frankle, J. and Carbin, M · 2019
Later among the works it cites.
Adversarial examples are a natural consequence of test error in noise
Gilmer, J., Ford, N., Carlini, N., and Cubuk, E · 2019
Later among the works it cites.
First-order adversarial vulnerability of neural networks and input dimension
Simon-Gabriel, C.-J., Ollivier, Y., Bottou, L., Schölkopf, B., and Lopez-Paz, D · 2019
Later among the works it cites.
Residual learning without normalization via better initialization
Zhang, H., Dauphin, Y. N., and Ma, T · 2019
Later among the works it cites.
Rethinking softmax cross-entropy loss for adversarial robustness
Pang, T., Xu, K., Dong, Y., Du, C., Chen, N., and Zhu, J · 2020
Closest in time.
Four things everyone should know to improve batch normalization
Summers, C. and Dinneen, M. J · 2020
Closest in time.
Intriguing properties of adversarial training at scale
Xie, C. and Yuille, A · 2020
Closest in time.