Fetching the paper…
Reading the bibliography…
Supervised learning methods trained with maximum likelihood objectives often overfit on training data.
On evaluating adversarial robustness
Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A. (2019) · 1902
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T. (2019) · 1903
Earlier work this paper cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Zhang, J., He, T., Sra, S., and Jadbabaie, A. (2019b) · 1905
Earlier work this paper cites.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G. (2019) · 1906
Earlier work this paper cites.
Adversarial training can hurt generalization
Raghunathan, A., Xie, S. M., Yang, F., Duchi, J. C., and Liang, P. (2019) · 1906
Earlier work this paper cites.
Your classifier is secretly an energy based model and you should treat it like one
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K. (2019) · 1912
Earlier work this paper cites.
Probability: The deductive and inductive problems
Johnson, W. E. (1932) · 1932
Earlier work this paper cites.
Information theory and statistical mechanics
Jaynes, E. T. (1957) · 1957
Earlier work this paper cites.
A comparison of the enhanced good-turing and deleted estimation methods for estimating probabilities of english bigrams
Church, K. W. and Gale, W. A. (1991) · 1991
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A. (1992) · 1992
Earlier work this paper cites.
Virtual adversarial training: a regularization method for supervised and semi-supervised learning
Miyato, T., Maeda, S.-i., Koyama, M., and Ishii, S. (2018) · 1993
Earlier work this paper cites.
Zero duality gap for a class of nonconvex optimization problems
Li, D. (1995) · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R. (1996) · 1996
Earlier work this paper cites.
Optimization of conditional value-at-risk
Rockafellar, R. T., Uryasev, S., et al. (2000) · 2000
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Koltchinskii, V. and Panchenko, D. (2002) · 2002
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Grandvalet, Y. and Bengio, Y. (2004) · 2004
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
Wang, D., Shelhamer, E., Liu, S., Olshausen, B., and Darrell, T. (2020) · 2006
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Kakade, S. M., Sridharan, K., and Tewari, A. (2008) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. et al. (2009) · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Robust variable selection with exponential squared loss
Wang, X., Jiang, Y., Huang, M., and Zhang, H. (2013) · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C. (2014) · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Le, Y. and Yang, X. (2015) · 2015
Cited alongside, same era.
Optimizing the cvar via sampling
Tamar, A., Glassner, Y., and Mannor, S. (2015) · 2015
Cited alongside, same era.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Adversarial machine learning at scale
Kurakin, A., Goodfellow, I., and Bengio, S. (2016) · 2016
Cited alongside, same era.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A. (2018) · 2018
Later among the works it cites.
Generalizing to unseen domains via adversarial data augmentation
Volpi, R., Namkoong, H., Sener, O., Duchi, J., Murino, V., and Savarese, S. (2018) · 2018
Later among the works it cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. (2018) · 2018
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q. (2019) · 2019
Later among the works it cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepfool: a simple and accurate method to fool deep neural networks
Moosavi-Dezfooli, S.-M., Fawzi, A., and Frossard, P. (2016) · 2016
Cited alongside, same era.
Semi-supervised learning with generative adversarial networks
Odena, A. (2016) · 2016
Cited alongside, same era.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N. (2016) · 2016
Cited alongside, same era.
Adversarial manipulation of deep representations
Sabour, S., Cao, Y., Faghri, F., and Fleet, D. J. (2016) · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016) · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Sgd on neural networks learns functions of increasing complexity
Kalimeris, D., Kaplun, G., Nakkiran, P., Edelman, B., Yang, T., Barak, B., and Zhang, H. (2019) · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z. (2019) · 2019
Later among the works it cites.
Semi-supervised domain adaptation via minimax entropy
Saito, K., Kim, D., Sclaroff, S., Darrell, T., and Saenko, K. (2019) · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Shorten, C. and Khoshgoftaar, T. M. (2019) · 2019
Later among the works it cites.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A. (2019) · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Wainwright, M. J. (2019) · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y. (2019) · 2019
Later among the works it cites.
Invariant risk minimization
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2020) · 2020
Later among the works it cites.
Towards a better understanding of label smoothing in neural machine translation
Gao, Y., Wang, W., Herold, C., Yang, Z., and Ney, H. (2020) · 2020
Later among the works it cites.
Understanding and mitigating the tradeoff between robustness and accuracy
Raghunathan, A., Xie, S. M., Yang, F., Duchi, J., and Liang, P. (2020) · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
Sagawa, S., Raghunathan, A., Koh, P. W., and Liang, P. (2020) · 2020
Later among the works it cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., and Ng, R. (2020) · 2020
Later among the works it cites.
Maximum-entropy adversarial data augmentation for improved generalization and robustness
Zhao, L., Liu, T., Peng, X., and Metaxas, D. (2020) · 2020
Later among the works it cites.
Mix-maxent: Improving accuracy and uncertainty estimates of deterministic neural networks
Pinto, F., Yang, H., Lim, S.-N., Torr, P., and Dokania, P. K. (2021) · 2021
Later among the works it cites.
Robustness and generalization via generative adversarial training
Poursaeed, O., Jiang, T., Yang, H., Belongie, S., and Lim, S.-N. (2021) · 2021
Later among the works it cites.
Salient imagenet: How to discover spurious features in deep learning?
Singla, S. and Feizi, S. (2021) · 2021
Later among the works it cites.
Training on test data with bayesian adaptation for covariate shift
Zhou, A. and Levine, S. (2021) · 2021
Later among the works it cites.