Fetching the paper…
Reading the bibliography…
Data augmentation is a widely used training trick in deep learning to improve the network generalization ability.
Inexact proximal-point penalty methods for non-convex optimization with non-convex constraints
Qihang Lin, Runchao Ma, and Yangyang Xu · 1908
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Document image defect models
Henry S Baird · 1992
Earlier work this paper cites.
On the convergence of the newton/log-barrier method
Stephen J Wright · 2001
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization : a basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Introduction to nonparametric estimation
Alexandre B Tsybakov · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
On the evaluation complexity of composite function minimization with applications to nonconvex nonlinear programming
Coralia Cartis, Nicholas IM Gould, and Philippe L Toint · 2011
Earlier work this paper cites.
Information theory: coding theorems for discrete memoryless systems
Imre Csiszar and János Körner · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Cited alongside, same era.
Faster autoaugment: Learning augmentation strategies using backpropagation
Ryuichiro Hataya, Jan Zdenek, Kazuki Yoshizoe, and Hideki Nakayama · 2019
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li · 2019
Later among the works it cites.
Population based augmentation: Efficient learning of augmentation policy schedules
Daniel Ho, Eric Liang, Xi Chen, Ion Stoica, and Pieter Abbeel · 2019
Later among the works it cites.
Fast autoaugment
Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and Sungwoong Kim · 2019
Later among the works it cites.
Proximally constrained methods for weakly convex optimization with weakly convex constraints
Runchao Ma, Qihang Lin, and Tianbao Yang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
A simple proximal stochastic gradient method for nonsmooth nonconvex optimization
Zhize Li and Jian Li · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Between-class learning for image classification
Yuji Tokozume, Yoshitaka Ushiku, and Tatsuya Harada · 2018
Cited alongside, same era.
A unified analysis of stochastic momentum methods for deep learning
Yan Yan, Tianbao Yang, Zhe Li, Qihang Lin, and Yi Yang · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Does data augmentation lead to positive margin?
Shashank Rajput, Zhili Feng, Zachary Charles, Po-Ling Loh, and Dimitris Papailiopoulos · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Later among the works it cites.
Spiderboost and momentum: Faster variance reduction algorithms
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2019
Later among the works it cites.
Stagewise training accelerates convergence of testing error over sgd
Zhuoning Yuan, Yan Yan, Rong Jin, and Tianbao Yang · 2019
Later among the works it cites.
Small relu networks are powerful memorizers: a tight analysis of memorization capacity
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
Later among the works it cites.
Learning data augmentation strategies for object detection
Barret Zoph, Ekin D Cubuk, Golnaz Ghiasi, Tsung-Yi Lin, Jonathon Shlens, and Quoc V Le · 2019
Later among the works it cites.
Complexity and performance of an augmented lagrangian algorithm
EG Birgin and JM Martínez · 2020
Closest in time.
More data can expand the generalization gap between adversarially robust and standard models
Lin Chen, Yifei Min, Mingrui Zhang, and Amin Karbasi · 2020
Closest in time.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Closest in time.
Exponential step sizes for non-convex optimization
Xiaoyu Li, Zhenxun Zhuang, and Francesco Orabona · 2020
Closest in time.
Yifei Min, Lin Chen, and Amin Karbasi · 2020
Closest in time.
A log-barrier newton-cg method for bound constrained optimization with complexity guarantees
Michael O’Neill and Stephen J Wright · 2020
Closest in time.
Understanding and mitigating the tradeoff between robustness and accuracy
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, and Percy Liang · 2020
Closest in time.
On the generalization effects of linear transformations in data augmentation
Sen Wu, Hongyang R Zhang, Gregory Valiant, and Christopher Ré · 2020
Closest in time.
Towards understanding label smoothing
Yi Xu, Yuanhong Xu, Qi Qian, Hao Li, and Rong Jin · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Closest in time.