Fetching the paper…
Reading the bibliography…
Training algorithms, broadly construed, are an essential part of every deep learning pipeline.
Some methods of speeding up the convergence of iteration methods
Boris T. Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
An Introduction to the Bootstrap
Bradley Efron and Robert J. Tibshirani · 1993
Earlier work this paper cites.
TIMIT Acoustic-Phonetic Continuous Speech Corpus
John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, David S. Pallett, Nancy L. Dahlgren, and Victor Zue · 1993
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen Wright · 1999
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Elizabeth D. Dolan and Jorge J. Moré · 2002
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
Alex Graves, Santiago Fernandez, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Reducing the Dimensionality of Data with Neural Networks
Geoffrey E. Hinton and Ruslan R. Salakhutdinov · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Findings of the 2014 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna · 2014
Earlier work this paper cites.
Criteo 1TB Click Logs dataset
Criteo A. I. Lab · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2014
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Findings of the 2015 Workshop on Statistical Machine Translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
James Martens and Roger B. Grosse · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Tony Han, Awni Y. Hannun, Billy Jun, Patrick LeGresley, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Yi Wang, Zhiqian Wang, Chong Wang, Bo Xiao, Dani Yogatama, Jun Zhan, and Zhenyao Zhu · 2016
Earlier work this paper cites.
Incorporating Nesterov Momentum into Adam
Timothy Dozat · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs), 2016
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Wide Residual Networks, 2016
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Findings of the 2017 Conference on Machine Translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi · 2017
Earlier work this paper cites.
Critical Hyper-Parameters: No Random, No Cry, 2017
Olivier Bousquet, Sylvain Gelly, Karol Kurach, Olivier Teytaud, and Damien Vincent · 2017
Earlier work this paper cites.
DAWNBench: An end-to-end deep learning benchmark and competition
Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia · 2017
Earlier work this paper cites.
Language Modeling with Gated Convolutional Networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Using Google Cloud Machine Learning to predict clicks at scale, 2017
Andreas Sterbenz · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Earlier work this paper cites.
Large Batch Training of Convolutional Networks, 2017
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks, 2018
Peter Battaglia, Jessica Blake Chandler Hamrick, Victor Bapst, Alvaro Sanchez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andy Ballard, Justin Gilmer, George E. Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Jayne Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu · 2018
Cited alongside, same era.
Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Cited alongside, same era.
Shampoo: Preconditioned Stochastic Tensor Optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Practical Quasi-Newton Methods for Training Deep Neural Networks
Yi Ren, Achraf Bahamou, and Donald Goldfarb · 2020
Later among the works it cites.
Optimizer Benchmarking Needs to Acccount for Hyperparameter Tuning
Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, and François Fleuret · 2020
Later among the works it cites.
On Layer Normalization in the Transformer Architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You, Jing Li, Shashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C. Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Call for Clarity in Reporting BLEU Scores
Matt Post · 2018
Cited alongside, same era.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
MoleculeNet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande · 2018
Cited alongside, same era.
fastMRI: An Open Dataset and Benchmarks for Accelerated MRI, 2018
Jure Zbontar, Florian Knoll, Anuroop Sriram, Tullie Murrell, Zhengnan Huang, Matthew J. Muckley, Aaron Defazio, Ruben Stern, Patricia Johnson, Mary Bruno, Marc Parente, Krzysztof J. Geras, Joe Katsnelson, Hersh Chandarana, Zizhao Zhang, Michal Drozdzal, Adriana Romero, Michael Rabbat, Pascal Vincent, Nafissa Yakubova, James Pinkerton, Duo Wang, Erich Owens, C. Lawrence Zitnick, Michael P. Recht, Daniel K. Sodickson, and Yvonne W. Lui · 2018
Cited alongside, same era.
mixup: Beyond Empirical Risk Minimization
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
TBD: Benchmarking and Analyzing Deep Neural Network Training
Hongyu Zhy, Mohamed Akrout, Bojian Zheng, Andrew Pelegris, Amar Phanishayee, Bianca Schroeder, and Gennady Pekhimenko · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Revisiting ResNets: Improved Training and Scaling Strategies
Irwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk, Aravind Srinivas, Tsung-Yi Lin, Jonathon Shlens, and Barret Zoph · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
A Loss Curvature Perspective on Training Instabilities of Deep Learning Models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George E. Dahl, Zachary Nado, and Orhan Firat · 2021
Later among the works it cites.
The Lipschitz Constant of Self-Attention
Hyunjik Kim, George Papamakarios, and Andriy Mnih · 2021
Later among the works it cites.
On the Adequacy of Untuned Warmup for Adaptive Optimization
Jerry Ma and Dennis Yarats · 2021
Later among the works it cites.
A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes, 2021
Zachary Nado, Justin M. Gilmer, Christopher J. Shallue, Rohan Anil, and George E. Dahl · 2021
Later among the works it cites.
Tensor normal training for deep learning models
Yi Ren and Donald Goldfarb · 2021
Later among the works it cites.
Descending through a Crowded Valley – Benchmarking Deep Learning Optimizers
Robin M. Schmidt, Frank Schneider, and Philipp Hennig · 2021
Later among the works it cites.
Technical Report for OGB Graph Property Prediction
Yan Wang, Hao Zhang, Jing Yang, Ruixin Zhang, and Shouhong Ding · 2021
Later among the works it cites.
ResNet strikes back: An improved training procedure in timm, 2021
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Greg Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Later among the works it cites.
LocoProp: Enhancing BackProp via Local loss optimization
Ehsan Amid, Rohan Anil, and Manfred Warmuth · 2022
Later among the works it cites.
Predicting the utility of search spaces for black-box optimization: a simple, budget-aware approach
Setareh Ariafar, Justin Gilmer, Zachary Nado, Jasper Snoek, Rodolphe Jenatton, and George E. Dahl · 2022
Later among the works it cites.
Amortized proximal optimization
Juhan Bae, Paul Vicol, Jeff Z. HaoChen, and Roger B. Grosse · 2022
Later among the works it cites.
Better plain ViT baselines for ImageNet-1k, 2022
Lucas Beyer, Xiaohua Zhai, and Alexander Kolesnikov · 2022
Later among the works it cites.
Adaptive Gradient Methods at the Edge of Stability, 2022
Jeremy M. Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E. Dahl, and Justin Gilmer · 2022
Later among the works it cites.
Bi-SimCut: A Simple Strategy for Boosting Neural Machine Translation
Pengzhi Gao, Zhongjun He, Hua Wu, and Haifeng Wang · 2022
Later among the works it cites.
Benchopt: Reproducible, efficient and collaborative optimization benchmarks
Thomas Moreau, Mathurin Massias, Alexandre Gramfort, Pierre Ablin, Pierre-Antoine Bannier, Benjamin Charlier, Mathieu Dagréou, Tom Dupré la Tour, Ghislain Durif, Cassio F. Dantas, Quentin Klopfenstein, Johan Larsson, En Lai, Tanguy Lefort, Benoit Malrézieux, Badr Moufad, Binh T. Nguyen, Alain Rakotomamonjy, Zaccharie Ramzi, Joseph Salmon, and Samuel Vaiter · 2022
Later among the works it cites.
How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer · 2022
Later among the works it cites.
Amos: An Adam-style Optimizer with Adaptive Weight Decay towards Model-Oriented Scale, 2022
Ran Tian and Ankur P. Parikh · 2022
Later among the works it cites.
Pooling Architecture Search for Graph Property Prediction in Open Graph Benchmark
Xu Wang, Huan Zhao, Lanning Wei, and Quanming Yao · 2022
Later among the works it cites.
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models, 2022
Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan · 2022
Later among the works it cites.
RePAST: A ReRAM-based PIM Accelerator for Second-order Training of DNN, 2022
Yilong Zhao, Li Jiang, Mingyu Gao, Naifeng Jing, Chengyang Gu, Qidong Tang, Fangxin Liu, Tao Yang, and Xiaoyao Liang · 2022
Later among the works it cites.
Symbolic Discovery of Optimization Algorithms, 2023
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
Scaling Vision Transformers to 22 Billion Parameters, 2023
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Patrick Collier, Alexey Gritsenko, Vighnesh Birodkar, Cristina Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov, Filip Pavetić, Dustin Tran, Thomas Kipf, Mario Lučić, Xiaohua Zhai, Daniel Keysers, Jeremiah Harmsen, and Neil Houlsby · 2023
Closest in time.
Benchmarking Graph Neural Networks
Vijay Prakash Dwivedi, Chaitanya K. Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson · 2023
Closest in time.
Deep Learning Tuning Playbook, 2023
Varun Godbole, George E. Dahl, Justin Gilmer, Christopher J. Shallue, and Zachary Nado · 2023
Closest in time.
Flax: A neural network library and ecosystem for JAX, 2023
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2023
Closest in time.
Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2023
Closest in time.
DLRM for PyTorch, 2023
NVIDIA · 2023
Closest in time.