Fetching the paper…
Reading the bibliography…
Large-batch training has become a commonly used technique when training neural networks with a large number of GPU/TPU processors.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Forrest N Iandola, Matthew W Moskewicz, Khalid Ashraf, and Kurt Keutzer · 2016
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow · 2016
Earlier work this paper cites.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Earlier work this paper cites.
Valeriu Codreanu, Damian Podareanu, and Vikram Saletore · 2017
Earlier work this paper cites.
Adabatch: Adaptive batch sizes for training deep neural networks
Aditya Devarakonda, Maxim Naumov, and Michael Garland · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Minimax statistical learning with wasserstein distances
Jaeho Lee and Maxim Raginsky · 2017
Earlier work this paper cites.
Scaling distributed machine learning with system and algorithm co-design
Mu Li · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Cited alongside, same era.
Certifiable distributional robustness with principled adversarial training
Aman Sinha, Hongseok Namkoong, and John Duchi · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Robustness via curvature regularization, and vice versa
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Jonathan Uesato, and Pascal Frossard · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Adversarial training for free!
Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein · 2019
Later among the works it cites.
On the convergence and robustness of adversarial training
Yisen Wang, Xingjun Ma, James Bailey, Jinfeng Yi, Bowen Zhou, and Quanquan Gu · 2019
Later among the works it cites.
Yet another accelerated sgd: Resnet-50 training on imagenet in 74.7 seconds
Masafumi Yamazaki, Akihiko Kasagi, Akihiro Tabuchi, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, and Kohta Nakashima · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Cited alongside, same era.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Cited alongside, same era.
Image classification at supercomputer scale
Chris Ying, Sameer Kumar, Dehao Chen, Tao Wang, and Youlong Cheng · 2018
Cited alongside, same era.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Cited alongside, same era.
Multigrain: a unified image embedding for classes and instances
Maxim Berman, Hervé Jégou, Andrea Vedaldi, Iasonas Kokkinos, and Matthijs Douze · 2019
Cited alongside, same era.
Autoaugment: Learning augmentation strategies from data
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Later among the works it cites.
Large-batch training for lstm and beyond
Yang You, Jonathan Hseu, Chris Ying, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
You only propagate once: Accelerating adversarial training via maximal principle
Dinghuai Zhang, Tianyuan Zhang, Yiping Lu, Zhanxing Zhu, and Bin Dong · 2019
Later among the works it cites.
Understanding and improving fast adversarial training
Maksym Andriushchenko and Nicolas Flammarion · 2020
Later among the works it cites.
Fast is better than free: Revisiting adversarial training
Eric Wong, Leslie Rice, and J Zico Kolter · 2020
Later among the works it cites.
Adversarial examples improve image recognition
Cihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang, Alan L Yuille, and Quoc V Le · 2020
Later among the works it cites.
Robust and accurate object detection via adversarial learning, 2021
Xiangning Chen, Cihang Xie, Mingxing Tan, Li Zhang, Cho-Jui Hsieh, and Boqing Gong · 2021
Closest in time.
Adversarial masking: Towards understanding robustness trade-off for generalization, 2021
Minhao Cheng, Zhe Gan, Yu Cheng, Shuohang Wang, Cho-Jui Hsieh, and Jingjing Liu · 2021
Closest in time.
Exploring the limits of concurrency in ml training on google tpus
Sameer Kumar, Yu Wang, Cliff Young, James Bradbury, Naveen Kumar, Dehao Chen, and Andy Swing · 2021
Closest in time.