Fetching the paper…
Reading the bibliography…
Regularization is a fundamental technique to prevent over-fitting and to improve generalization performances by constraining a model's complexity.
Exploring bias in gan-based data augmentation for small samples
Hu, M. and Li, J. (2019) · 1905
Earlier work this paper cites.
Implicit rugosity regularization via data augmentation
LeJeune, D., Balestriero, R., Javadi, H., and Baraniuk, R. G. (2019) · 1905
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Ix. on the problem of the most efficient tests of statistical hypotheses
Neyman, J. and Pearson, E. S. (1933) · 1933
Earlier work this paper cites.
On the stability of inverse problems
Tikhonov, A. N. (1943) · 1943
Earlier work this paper cites.
The generalization of ‘student’s’problem when several different population varlances are involved
Welch, B. L. (1947) · 1947
Earlier work this paper cites.
Statistical methods and scientific induction
Fisher, R. (1955) · 1955
Earlier work this paper cites.
Solution of incorrectly formulated problems and the regularization method
Tihonov, A. N. (1963) · 1963
Earlier work this paper cites.
The method of ordered risk minimization, i
Vapnik, V. and Chervonenkis, A. Y. (1974) · 1974
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. (1991) · 1991
Earlier work this paper cites.
Tangent prop-a formalism for specifying selected invariances in an adaptive network
Simard, P., Victorri, B., LeCun, Y., and Denker, J. (1991) · 1991
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Kohavi, R., Wolpert, D. H., et al. (1996) · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Vicinal risk minimization
Chapelle, O., Weston, J., Bottou, L., and Vapnik, V. (2000) · 2000
Earlier work this paper cites.
Understanding and mitigating the tradeoff between robustness and accuracy
Raghunathan, A., Xie, S. M., Yang, F., Duchi, J., and Liang, P. (2020) · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K. (2020b) · 2003
Earlier work this paper cites.
Bias-variance analysis of support vector machines for the development of svm-based ensemble methods
Valentini, G. and Dietterich, T. G. (2004) · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. and Nasrabadi, N. M. (2006) · 2006
Earlier work this paper cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Hui, L. and Belkin, M. (2020) · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. (2009) · 2009
Cited alongside, same era.
Efficient and accurate lp-norm multiple kernel learning
Kloft, M., Brefeld, U., Laskov, P., Müller, K.-R., Zien, A., and Sonnenburg, S. (2009) · 2009
Cited alongside, same era.
A survey on transfer learning
Pan, S. J. and Yang, Q. (2009) · 2009
Cited alongside, same era.
Wemix: How to better utilize data augmentation
Xu, Y., Noy, A., Lin, M., Qian, Q., Li, H., and Jin, R. (2020) · 2010
Cited alongside, same era.
Bayesian inference in statistical analysis
Box, G. E. and Tiao, G. C. (2011) · 2011
Cited alongside, same era.
Statistical learning theory: Models, concepts, and results
Von Luxburg, U. and Schölkopf, B. (2011) · 2011
Condensenet: An efficient densenet using learned group convolutions
Huang, G., Liu, S., Van der Maaten, L., and Weinberger, K. Q. (2018) · 2018
Later among the works it cites.
Dealing with bias via data augmentation in supervised learning scenarios
Iosifidis, V. and Ntoutsi, E. (2018) · 2018
Later among the works it cites.
Improving deep learning with generic data augmentation
Taylor, L. and Nitschke, G. (2018) · 2018
Later among the works it cites.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A. (2018) · 2018
Later among the works it cites.
The inaturalist species classification and detection dataset
Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic gradient descent tricks
Bottou, L. (2012) · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Cited alongside, same era.
Constrained optimization and Lagrange multiplier methods
Bertsekas, D. P. (2014) · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2014) · 2014
Cited alongside, same era.
Data augmentation for deep neural network acoustic modeling
Cui, X., Goel, V., and Kingsbury, B. (2015) · 2015
Cited alongside, same era.
Machine learning: Trends, perspectives, and prospects
Jordan, M. I. and Mitchell, T. M. (2015) · 2015
Cited alongside, same era.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Shorten, C. and Khoshgoftaar, T. M. (2019) · 2019
Later among the works it cites.
Manifold mixup: Better representations by interpolating hidden states
Verma, V., Lamb, A., Beckham, C., Najafi, A., Mitliagkas, I., Lopez-Paz, D., and Bengio, Y. (2019) · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y. (2019) · 2019
Later among the works it cites.
Theoretically principled trade-off between robustness and accuracy
Zhang, H., Yu, Y., Jiao, J., Xing, E., El Ghaoui, L., and Jordan, M. (2019) · 2019
Later among the works it cites.
Deflating dataset bias using synthetic data augmentation
Jaipuria, N., Zhang, X., Bhasin, R., Arafa, M., Chakravarty, P., Shrivastava, S., Manglani, S., and Murali, V. N. (2020) · 2020
Later among the works it cites.
Self-supervised learning of pretext-invariant representations
Misra, I. and Maaten, L. v. d. (2020) · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Xie, Q., Dai, Z., Hovy, E., Luong, T., and Le, Q. (2020) · 2020
Later among the works it cites.
Selecting data augmentation for simulating interventions
Ilse, M., Tomczak, J. M., and Forré, P. (2021) · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021) · 2021
Later among the works it cites.
The curious case of adversarially robust models: More data can help, double descend, or hurt generalization
Min, Y., Chen, L., and Karbasi, A. (2021) · 2021
Later among the works it cites.
Efficientnetv2: Smaller models and faster training
Tan, M. and Le, Q. (2021) · 2021
Later among the works it cites.
Balestriero, R., Misra, I., and LeCun, Y. (2022) · 2022
Closest in time.
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022) · 2022
Closest in time.