Fetching the paper…
Reading the bibliography…
The field of deep learning has witnessed significant progress, particularly in computer vision (CV), natural language processing (NLP), and speech.
Augmix: A simple data processing method to improve robustness and uncertainty
Dan Hendrycks, Norman Mu, Ekin D Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan · 1912
Earlier work this paper cites.
Xinyu Zhang, Qiang Wang, Jian Zhang, and Zhao Zhong · 1912
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Iterative procedures for nonlinear integral equations
Donald G Anderson · 1965
Earlier work this paper cites.
Quasi-newton methods, motivation and theory
John E Dennis, Jr and Jorge J Moré · 1977
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Jeffrey L Elman · 1993
Earlier work this paper cites.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard HR Hahnloser, Rahul Sarpeshkar, Misha A Mahowald, Rodney J Douglas, and H Sebastian Seung · 2000
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2006
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Shallow-to-deep training for neural machine translation
Bei Li, Ziyang Wang, Hui Liu, Yufan Jiang, Quan Du, Tong Xiao, Huizhen Wang, and Jingbo Zhu · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Anderson acceleration for fixed-point iterations
Homer F Walker and Peng Ni · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Random walk initialization for training very deep feedforward networks
David Sussillo and LF Abbott · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Butterfly factorization
Yingzhou Li, Haizhao Yang, Eileen R Martin, Kenneth L Ho, and Lexing Ying · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Dmytro Mishkin and Jiri Matas · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Convolutional neural networks with low-rank regularization
Cheng Tai, Tong Xiao, Yi Zhang, Xiaogang Wang, et al · 2015
Earlier work this paper cites.
That’s so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using# petpeeve tweets
William Yang Wang and Diyi Yang · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Regularized nonlinear acceleration
Damien Scieur, Alexandre d’Aspremont, and Francis Bach · 2016
Earlier work this paper cites.
Gradual dropin of layers to train very deep neural networks
Leslie N Smith, Emily M Hand, and Timothy Doster · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Freezeout: Accelerate training by progressively freezing layers
Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston · 2017
Earlier work this paper cites.
Improving deep learning by inverse square root linear units (isrlus)
Brad Carlile, Guy Delamarter, Paul Kinney, Akiko Marti, and Brian Whitney · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Biased importance sampling for deep neural network training
Angelos Katharopoulos and François Fleuret · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
On weight initialization in deep neural networks
Siddharth Krishna Kumar · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Nonlinear acceleration of stochastic algorithms
Damien Scieur, Francis Bach, and Alexandre d’Aspremont · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander Alemi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep growing learning
Guangcong Wang, Xiaohua Xie, Jianhuang Lai, and Jiaxuan Zhuo · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
On the computational inefficiency of large batch sizes for stochastic gradient descent
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W Mahoney, and Joseph Gonzalez · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Earlier work this paper cites.
Fastai - progressive resizing
Jeremy Howard · 2018
Earlier work this paper cites.
Decoupled parallel backpropagation with convergence guarantee
Zhouyuan Huo, Bin Gu, Heng Huang, et al · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and François Fleuret · 2018
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso · 2018
Earlier work this paper cites.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Earlier work this paper cites.
A faster algorithm for reducing the computational complexity of convolutional neural networks
Yulin Zhao, Donghui Wang, Leiou Wang, and Peng Liu · 2018
Earlier work this paper cites.
Metainit: Initializing learning by learning to initialize
Yann N Dauphin and Samuel Schoenholz · 2019
Earlier work this paper cites.
Efficient training of bert by progressively stacking
Linyuan Gong, Di He, Zhuohan Li, Tao Qin, Liwei Wang, and Tieyan Liu · 2019
Earlier work this paper cites.
Bag of tricks for image classification with convolutional neural networks
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li · 2019
Earlier work this paper cites.
Population based augmentation: Efficient learning of augmentation policy schedules
Daniel Ho, Eric Liang, Xi Chen, Ion Stoica, and Pieter Abbeel · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Accelerating deep learning by focusing on the biggest losers
Angela H Jiang, Daniel L-K Wong, Giulio Zhou, David G Andersen, Jeffrey Dean, Gregory R Ganger, Gauri Joshi, Michael Kaminksy, Michael Kozuch, Zachary C Lipton, et al · 2019
Cited alongside, same era.
Budgeted training: Rethinking deep neural network training under resource constraints
Mengtian Li, Ersin Yumer, and Deva Ramanan · 2019
Cited alongside, same era.
Fast autoaugment
Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and Sungwoong Kim · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Prunetrain: fast neural network training by dynamic sparse model reconfiguration
Sangkug Lym, Esha Choukse, Siavash Zangeneh, Wei Wen, Sujay Sanghavi, and Mattan Erez · 2019
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer · 2021
Later among the works it cites.
Efficientnetv2: Smaller models and faster training
Mingxing Tan and Quoc Le · 2021
Later among the works it cites.
1-bit adam: Communication efficient large-scale training with adam’s convergence speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Paddlepaddle: An open-source deep learning platform from industrial practice
Yanjun Ma, Dianhai Yu, Tian Wu, and Haifeng Wang · 2019
Cited alongside, same era.
Efficient winograd or cook-toom convolution kernel implementation on widely used mobile cpus
Partha Maji, Andrew Mundy, Ganesh Dasika, Jesse Beu, Matthew Mattina, and Robert Mullins · 2019
Cited alongside, same era.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Cited alongside, same era.
Pipedream: Generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
composer
The Mosaic ML Team · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX
Ben Wang · 2021
Later among the works it cites.
LightSeq: A high performance inference library for transformers
Xiaohui Wang, Ying Xiong, Yang Wei, Mingxuan Wang, and Lei Li · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Later among the works it cites.
Zero initialization: Initializing neural networks with only zeros and ones
Jiawei Zhao, Florian Schäfer, and Anima Anandkumar · 2021
Later among the works it cites.
Gradinit: Learning to initialize neural networks for stable and efficient training
Chen Zhu, Renkun Ni, Zheng Xu, Kezhi Kong, W Ronny Huang, and Tom Goldstein · 2021
Later among the works it cites.
Towards understanding sharpness-aware minimization
Maksym Andriushchenko and Nicolas Flammarion · 2022
Later among the works it cites.
Quasi-newton methods for machine learning: forget the past, just sample
Albert S Berahas, Majid Jahani, Peter Richtárik, and Martin Takáč · 2022
Later among the works it cites.
GPT-NeoX-20B: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach · 2022
Later among the works it cites.
Token merging: Your vit but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2022
Later among the works it cites.
Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition
Han Cai, Chuang Gan, and Song Han · 2022
Later among the works it cites.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Later among the works it cites.
A scalable pipeline for gigapixel whole slide imaging analysis on leadership class hpc systems
Sajal Dash, Benjamín Hernández, Aristeidis Tsaris, Folami T Alamudun, Hong-Jun Yoon, and Feivi Wang · 2022
Later among the works it cites.
Scenic: A jax library for computer vision research and beyond
Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab, Matthias Minderer, and Yi Tay · 2022
Later among the works it cites.
Sequential normalization: Embracing smaller sample sizes for normalization
Neofytos Dimitriou and Ognjen Arandjelović · 2022
Later among the works it cites.
Re-parameterizing your optimizers rather than architectures
Xiaohan Ding, Honghao Chen, Xiangyu Zhang, Kaiqi Huang, Jungong Han, and Guiguang Ding · 2022
Later among the works it cites.
Resist: Layer-wise decomposition of resnets for distributed training
Chen Dun, Cameron R Wolfe, Christopher M Jermaine, and Anastasios Kyrillidis · 2022
Later among the works it cites.
Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design
Hongxiang Fan, Thomas Chau, Stylianos I Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D Lane, and Mohamed S Abdelfattah · 2022
Later among the works it cites.
(beta) channels last memory format in pytorch, 2022
Vitaly Fedyunin · 2022
Later among the works it cites.
Cramming: Training a language model on a single gpu in one day
Jonas Geiping and Tom Goldstein · 2022
Later among the works it cites.
A survey on efficient convolutional neural networks and hardware acceleration
Deepak Ghimire, Dayoung Kil, and Seong-heum Kim · 2022
Later among the works it cites.
Turbo training with token dropout
Tengda Han, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Accelerating parallel stochastic gradient descent via non-blocking mini-batches
Haoze He and Parijat Dube · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le · 2022
Later among the works it cites.
Hardware-aware quantization/mapping strategies for compute-in-memory accelerators
Shanshi Huang, Hongwu Jiang, and Shimeng Yu · 2022
Later among the works it cites.
A data-loader tunable knob to shorten gpu idleness for distributed deep learning
Danlin Jia, Geng Yuan, Xue Lin, and Ningfang Mi · 2022
Later among the works it cites.
Stop wasting my time! saving days of imagenet and BERT training with latest weight averaging
Jean Kaddour · 2022
Later among the works it cites.
Zhenglun Kong, Haoyu Ma, Geng Yuan, Mengshu Sun, Yanyue Xie, Peiyan Dong, Xin Meng, Xuan Shen, Hao Tang, Minghai Qin, et al · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al · 2022
Later among the works it cites.
Ffcv: Accelerating training by removing data bottlenecks
Guillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park, Hadi Salman, and Aleksander Madry · 2022
Later among the works it cites.
xformers: A modular and hackable transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, and Daniel Haziza · 2022
Later among the works it cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2022
Later among the works it cites.
Tokenmix: Rethinking image mixing for data augmentation in vision transformers
Jihao Liu, Boxiao Liu, Hang Zhou, Hongsheng Li, and Yu Liu · 2022
Later among the works it cites.
Communication-efficient distributed learning for large batch optimization
Rui Liu and Barzan Mozafari · 2022
Later among the works it cites.
Maximizing communication efficiency for large-scale training via 0/1 adam
Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, and Yuxiong He · 2022
Later among the works it cites.
TorchScale: Transformers at scale
Shuming Ma, Hongyu Wang, Shaohan Huang, Wenhui Wang, Zewen Chi, Li Dong, Alon Benhaim, Barun Patra, Vishrav Chaudhary, Xia Song, and Furu Wei · 2022
Later among the works it cites.
Butterflyflow: Building invertible layers with butterfly matrices
Chenlin Meng, Linqi Zhou, Kristy Choi, Tri Dao, and Stefano Ermon · 2022
Later among the works it cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Sören Mindermann, Jan M Brauner, Muhammed T Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N Gomez, Adrien Morisot, Sebastian Farquhar, et al · 2022
Later among the works it cites.
Improving transformers with probabilistic attention keys
Tam Minh Nguyen, Tan Minh Nguyen, Dung DD Le, Duy Khuong Nguyen, Viet-Anh Tran, Richard Baraniuk, Nhat Ho, and Stanley Osher · 2022
Later among the works it cites.
Machine Learning Systems: Design and Implementation
Open Machine Learning Systems Community · 2022
Later among the works it cites.
A survey on textual entailment based question answering
Aarthi Paramasivam and S Jaya Nirmala · 2022
Later among the works it cites.
The devil in linear transformer
Zhen Qin, XiaoDong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong · 2022
Later among the works it cites.
Bitblade: Energy-efficient variable bit-precision hardware accelerator for quantized neural networks
Sungju Ryu, Hyungjun Kim, Wooseok Yi, Eunhwan Kim, Yulhwa Kim, Taesu Kim, and Jae-Joon Kim · 2022
Later among the works it cites.
What language model to train if you have one million gpu hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Bideman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al · 2022
Later among the works it cites.
Feature wise normalization: An effective way of normalizing data
Dalwinder Singh and Birmohan Singh · 2022
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Later among the works it cites.
Dnn training acceleration via exploring gpgpu friendly sparsity
Zhuoran Song, Yihong Xu, Han Li, Naifeng Jing, Xiaoyao Liang, and Li Jiang · 2022
Later among the works it cites.
Prioritizing samples in reinforcement learning with reducible loss
Shivakanth Sujit, Somjit Nath, Pedro HM Braga, and Samira Ebrahimi Kahou · 2022
Later among the works it cites.
Profiling and improving the pytorch dataloader for high-latency storage: A technical report
Ivan Svogor, Christian Eichenberger, Markus Spanring, Moritz Neun, and Michael Kopp · 2022
Later among the works it cites.
Adversarial attack and defense strategies of speaker recognition systems: A survey
Hao Tan, Le Wang, Huan Zhang, Junjian Zhang, Muhammad Shafiq, and Zhaoquan Gu · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Later among the works it cites.
What’s in my ai? a comprehensive analysis of datasets used to train gpt-1, gpt-2, gpt-3, gpt-neox-20b, megatron-11b, mt-nlg, and gopher, 2022
Alan D. Thompson · 2022
Later among the works it cites.
Efficient methods for natural language processing: a survey
Marcos Treviso, Tianchu Ji, Ji-Ung Lee, Betty van Aken, Qingqing Cao, Manuel R Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Pedro H Martins, et al · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Cross-domain collaborative normalization via structural knowledge
Haifeng Xia and Zhengming Ding · 2022
Later among the works it cites.
Large-batch optimization for dense visual predictions
Zeyue Xue, Jianming Liang, Guanglu Song, Zhuofan Zong, Liang Chen, Yu Liu, and Ping Luo · 2022
Later among the works it cites.
Zhewei Yao, Xiaoxia Wu, Conglong Li, Connor Holmes, Minjia Zhang, Cheng Li, and Yuxiong He · 2022
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments
Binhang Yuan, Yongjun He, Jared Quincy Davis, Tianyi Zhang, Tri Dao, Beidi Chen, Percy Liang, Christopher Re, and Ce Zhang · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Mixhead: Breaking the low-rank bottleneck in multi-head attention language models
Zhong Zhang, Nian Shao, Chongming Gao, Rui Miao, Qinli Yang, and Junming Shao · 2022
Later among the works it cites.
Avalanche: A pytorch library for deep continual learning
Antonio Carta, Lorenzo Pellegrini, Andrea Cossu, Hamed Hemati, and Vincenzo Lomonaco · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
Chataug: Leveraging chatgpt for text data augmentation
Haixing Dai, Zhengliang Liu, Wenxiong Liao, Xiaoke Huang, Zihao Wu, Lin Zhao, Wei Liu, Ninghao Liu, Sheng Li, Dajiang Zhu, et al · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Closest in time.
Flax: A neural network library and ecosystem for JAX, 2023
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2023
Closest in time.
Accelerating self-supervised learning via efficient training strategies
Mustafa Taha Koçyiğit, Timothy M Hospedales, and Hakan Bilen · 2023
Closest in time.
HRBP: Hardware-friendly regrouping towards block-wise pruning for sparse training, 2023
Haoyu Ma, Chengming Zhang, lizhi xiang, Xiaolong Ma, Geng Yuan, Wenkai Zhang, Shiwei Liu, Tianlong Chen, Dingwen Tao, Yanzhi Wang, Zhangyang Wang, and Xiaohui Xie · 2023
Closest in time.
Simpletrack: Understanding and rethinking 3d multi-object tracking
Ziqi Pang, Zhichao Li, and Naiyan Wang · 2023
Closest in time.
Speechain: A speech toolkit for large-scale machine speech chain
Heli Qi, Sashi Novitasari, Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2023
Closest in time.
Ray: A distributed framework for emerging ai applications
Ray Project · 2023
Closest in time.
Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension
Anna Rogers, Matt Gardner, and Isabelle Augenstein · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
The transformer family version 2.0
Lilian Weng · 2023
Closest in time.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al · 2023
Closest in time.
A survey on efficient training of transformers
Bohan Zhuang, Jing Liu, Zizheng Pan, Haoyu He, Yuetian Weng, and Chunhua Shen · 2023
Closest in time.