Fetching the paper…
Reading the bibliography…
In an increasing number of domains it has been demonstrated that deep learning models can be trained using relatively large batch sizes without sacrificing data efficiency.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Trueskill™: A bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Richard H. Byrd, Gillian M. Chin, Jorge Nocedal, and Yuchen Wu · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling, 2013, 1312.3005
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
Auto-encoding variational bayes, 2013, arXiv:1312.6114
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems, 2014, 1412.6544
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014, 1412.6980
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks, 2015, 1511.06807
Arvind Neelakantan, Luke Vilnis, Quoc V. Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Optimization methods for large-scale machine learning, 2016, 1606.04838
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Coupling adaptive batch sizes with learning rates, 2016, 1612.05086
Lukas Balles, Javier Romero, and Philipp Hennig · 2016
Earlier work this paper cites.
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Big batch sgd: Automated inference using adaptive batch sizes, 2016, 1610.05792
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima, 2016, 1609.04836
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Large batch training of convolutional networks, 2017, 1708.03888
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Later among the works it cites.
Gradient diversity: a key ingredient for scalable distributed learning, 2017, 1706.05699
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2017
Later among the works it cites.
Imagenet training in minutes, 2017, 1709.05011
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2017
Later among the works it cites.
Igor Adamski, Robert Adamski, Tomasz Grel, Adam Jędrych, Kamil Kaczmarek, and Henryk Michalewski · 2018
Closest in time.
AI and Compute, May 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Cited alongside, same era.
Distributed second-order optimization using kronecker-factored approximations, 2017
Jimmy Ba, Roger Grosse, and James Martens · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
Adabatch: Adaptive batch sizes for training deep neural networks, 2017, 1712.02029
Aditya Devarakonda, Maxim Naumov, and Michael Garland · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2017, 1706.02677
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Why momentum really works
Gabriel Goh · 2017
Cited alongside, same era.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically, 2017, 1712.00409
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Dario Amodei and Danny Hernandez · 2018
Closest in time.
OpenAI Five, Jun 2018
Greg Brockman, Brooke Chan, Przemyslaw Debiak, Christy Dennison, David Farhi, Rafal Józefowicz, Jakub Pachocki, Michael Petrov, Henrique Pondé, Jonathan Raiman, Szymon Sidor, Jie Tang, Filip Wolski, and Susan Zhang · 2018
Closest in time.
Large scale gan training for high fidelity natural image synthesis, 2018, 1809.11096
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Closest in time.
The effect of network width on the performance of large-batch training, 2018, 1806.03791
Lingjiao Chen, Hongyi Wang, Jinman Zhao, Dimitris Papailiopoulos, and Paraschos Koutris · 2018
Closest in time.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Closest in time.
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W. Mahoney, and Joseph Gonzalez · 2018
Closest in time.
Distributed prioritized experience replay, 2018, 1803.00933
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Closest in time.
Are deep policy gradient algorithms truly policy gradient algorithms?, 2018, 1811.02553
Andrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2018
Closest in time.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, Tiegang Chen, Guangxiao Hu, Shaohuai Shi, and Xiaowen Chu · 2018
Closest in time.
Measuring the intrinsic dimension of objective landscapes, 2018, 1804.08838
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Closest in time.
Scaling neural machine translation, 2018, 1806.00187
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Closest in time.
Large scale language modeling: Converging on 40gb of text in four hours, 2018, 1808.01371
Raul Puri, Robert Kirby, Nikolai Yakovenko, and Bryan Catanzaro · 2018
Closest in time.
Accelerated methods for deep reinforcement learning, 2018, 1803.02811
Adam Stooke and Pieter Abbeel · 2018
Closest in time.
Measuring the effects of data parallelism on neural network training, 2018, arXiv:1811.03600
Christopher J. Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2018
Closest in time.
Understanding short-horizon bias in stochastic meta-optimization, 2018, 1803.02021
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Closest in time.