Fetching the paper…
Reading the bibliography…
Mixup, which creates synthetic training instances by linearly interpolating random sample pairs, is a simple and yet effective regularization technique to boost the performance of deep models trained with SGD.
Augmenting data with mixup for sentence classification: An empirical study
Hongyu Guo, Yongyi Mao, and Richong Zhang · 1905
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
Vladimir A Marčenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
The dynamics of learning: A random matrix approach
Zhenyu Liao and Romain Couillet · 2018
Earlier work this paper cites.
Do cifar-10 classifiers generalize to cifar-10?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Earlier work this paper cites.
Mixup as directional adversarial training
Guillaume P Archambault, Yongyi Mao, Hongyu Guo, and Richong Zhang · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Aditya Sharad Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
On mixup training: Improved calibration and predictive uncertainty for deep neural networks
Sunil Thulasidasan, Gopinath Chennupati, Jeff A Bilmes, Tanmoy Bhattacharya, and Sarah Michalak · 2019
Cited alongside, same era.
Manifold mixup: Better representations by interpolating hidden states
Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani, Andrew M Saxe, and Haim Sompolinsky · 2020
Cited alongside, same era.
k-mixup regularization for deep learning via optimal transport, 2021
Kristjan Greenewald, Anming Gu, Mikhail Yurochkin, Justin Solomon, and Edward Chien · 2021
Later among the works it cites.
Early stopping in deep networks: Double descent and how to eliminate it
Reinhard Heckel and Fatih Furkan Yilmaz · 2021
Later among the works it cites.
When and how epochwise double descent happens
Cory Stephenson and Tyler Lee · 2021
Later among the works it cites.
How does mixup help with robustness and generalization?
Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani, and James Zou · 2021
Later among the works it cites.
Label noise in adversarial training: A novel perspective to study robust overfitting
Chengyu Dong, Liyuan Liu, and Jingbo Shang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization of two-layer neural networks: An asymptotic viewpoint
Jimmy Ba, Murat Erdogdu, Taiji Suzuki, Denny Wu, and Tianzong Zhang · 2020
Cited alongside, same era.
Puzzle mix: Exploiting saliency and local statistics for optimal mixup
Jang-Hyun Kim, Wonho Choo, and Hyun Oh Song · 2020
Cited alongside, same era.
Early-learning regularization prevents memorization of noisy labels
Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda · 2020
Cited alongside, same era.
Harder or different? a closer look at distribution shift in dataset reproduction
Shangyun Lu, Bradley Nott, Aaron Olson, Alberto Todeschini, Hossein Vahabi, Yair Carmon, and Ludwig Schmidt · 2020
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Cited alongside, same era.
Overfitting in adversarially robust deep learning
Leslie Rice, Eric Wong, and Zico Kolter · 2020
Cited alongside, same era.
Rethinking bias-variance trade-off for generalization of neural networks
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma · 2020
Cited alongside, same era.
Xiaotian Han, Zhimeng Jiang, Ninghao Liu, and Xia Hu · 2022
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2022
Later among the works it cites.
Robust training under label noise by over-parameterization
Sheng Liu, Zhihui Zhu, Qing Qu, and Chong You · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Multi-scale feature learning dynamics: Insights for double descent
Mohammad Pezeshki, Amartya Mitra, Yoshua Bengio, and Guillaume Lajoie · 2022
Later among the works it cites.
Regmixup: Mixup as a regularizer can surprisingly improve accuracy and out distribution robustness
Francesco Pinto, Harry Yang, Ser-Nam Lim, Philip HS Torr, and Puneet K Dokania · 2022
Later among the works it cites.
Genlabel: Mixup relabeling using generative models
Jy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon, Dimitris Papailiopoulos, and Kangwook Lee · 2022
Later among the works it cites.
On the generalization of models trained with SGD: Information-theoretic bounds and implications
Ziqiao Wang and Yongyi Mao · 2022
Later among the works it cites.