Fetching the paper…
Reading the bibliography…
Recent advances in Transformers have come with a huge requirement on computing resources, highlighting the importance of developing efficient training techniques to make Transformer training faster, at lower cost, and to higher accuracy by the efficient use of computation and memory resources.
A method for unconstrained convex minimization problem with the rate of convergence o (1/k
Y. Nesterov · 1983
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
T. Chen, B. Xu, C. Zhang, and C. Guestrin · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, et al · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
N. P. Jouppi, C. Young, N. Patil, D. Patterson, et al · 2017
Earlier work this paper cites.
Convergence analysis of proximal gradient with momentum for nonconvex optimization
Q. Li, Y. Zhou, Y. Liang, and P. K. Varshney · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Y. You, I. Gitman, and B. Ginsburg · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Earlier work this paper cites.
Mixed precision training of convolutional neural networks using integer operations
D. Das, N. Mellempudi, D. Mudigere, D. D. Kalamkar, et al · 2018
Earlier work this paper cites.
Training deep models faster with robust, approximate importance sampling
T. B. Johnson and C. Guestrin · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
A. Katharopoulos and F. Fleuret · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Earlier work this paper cites.
Mixed precision training
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, et al · 2018
Earlier work this paper cites.
Outerspace: An outer product based sparse matrix multiplication accelerator
S. Pal, J. Beaumont, D.-H. Park, A. Amarnath, S. Feng, et al · 2018
Earlier work this paper cites.
Training deep neural networks with 8-bit floating point numbers
N. Wang, J. Choi, D. Brand, C.-Y. Chen, and K. Gopalakrishnan · 2018
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Earlier work this paper cites.
Efficient training of bert by progressively stacking
L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu · 2019
Earlier work this paper cites.
J. Herrmann, O. Beaumont, L. Eyraud-Dubois, J. Hermann, A. Joly, and A. Shilova · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, et al · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. D. M.-W. C. Kenton and L. K. Toutanova · 2019
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
N. Lee, T. Ajanthan, and P. Torr · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
The evolved transformer
D. So, Q. Le, and C. Liang · 2019
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
E. Strubell, A. Ganesh, and A. McCallum · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2019
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
H. Zhang, Y. N. Dauphin, and T. Ma · 2019
Cited alongside, same era.
Why adam beats sgd for attention models
J. Zhang, S. P. Karimireddy, A. Veit, S. Kim, et al · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Shifted and squeezed 8-bit floating point format for low-precision training of deep neural networks
L. Cambier, A. Bhiwandiwalla, T. Gong, M. Nekuii, O. H. Elibol, and H. Tang · 2020
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Later among the works it cites.
Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang · 2021
Later among the works it cites.
Accelerating sparse deep neural networks
A. Mishra, J. A. Latorre, J. Pool, D. Stosic, D. Stosic, et al · 2021
Later among the works it cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, et al · 2021
Later among the works it cites.
Mesa: A memory-saving training framework for transformers
Z. Pan, P. Chen, H. He, J. Liu, J. Cai, and B. Zhuang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J.-B. Cordonnier, A. Loukas, and M. Jaggi · 2020
Cited alongside, same era.
Batch normalization biases residual blocks towards the identity function in deep networks
S. De and S. Smith · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen · 2020
Cited alongside, same era.
Improving transformer optimization through better initialization
X. S. Huang, F. Perez, J. Ba, and M. Volkovs · 2020
Cited alongside, same era.
torchgpipe: On-the-fly pipeline parallelism for training giant models
C. Kim, H. Lee, M. Jeong, W. Baek, B. Yoon, et al · 2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Cited alongside, same era.
Train big, then compress: Rethinking model size for efficient training and inference of transformers
Z. Li, E. Wallace, S. Shen, K. Lin, et al · 2020
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
M. Paul, S. Ganguli, and G. K. Dziugaite · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, et al · 2021
Later among the works it cites.
Zero-offload: Democratizing billion-scale model training
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, et al · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Going deeper with image transformers
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou · 2021
Later among the works it cites.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2022
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations
X. Chen, C.-J. Hsieh, and B. Gong · 2022
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Later among the works it cites.
Efficient sharpness-aware minimization for improved training of neural networks
J. Du, H. Yan, J. Feng, J. T. Zhou, L. Zhen, et al · 2022
Later among the works it cites.
Sharpness-aware training for free
J. Du, D. Zhou, J. Feng, V. Y. Tan, and J. T. Zhou · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Later among the works it cites.
Visual prompt tuning
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim · 2022
Later among the works it cites.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
J. Kaddour · 2022
Later among the works it cites.
Automated progressive learning for efficient training of vision transformers
C. Li, B. Zhuang, G. Wang, X. Liang, X. Chang, and Y. Yang · 2022
Later among the works it cites.
Scaling language-image pre-training via masking
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He · 2022
Later among the works it cites.
GACT: activation compressed training for generic network architectures
X. Liu, L. Zheng, D. Wang, Y. Cen, W. Chen, X. Han, J. Chen, et al · 2022
Later among the works it cites.
Dota: detect and omit weak attentions for scalable transformer acceleration
Z. Qu, L. Liu, F. Tu, Z. Chen, Y. Ding, and Y. Xie · 2022
Later among the works it cites.
S. Smith, M. Patwary, B. Norick, et al · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. S. Morcos · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
E. B. Zaken, Y. Goldberg, and S. Ravfogel · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, et al · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, et al · 2023
Closest in time.
Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design
H. You, Z. Sun, H. Shi, Z. Yu, Y. Zhao, Y. Zhang, C. Li, B. Li, and Y. Lin · 2023
Closest in time.