Fetching the paper…
Reading the bibliography…
The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes.
Schwartz, R.; Dodge, J.; Smith, N. A.; and Etzioni, O. 2019 · 1907
Earlier work this paper cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Shoeybi, M.; Patwary, M.; Puri, R.; LeGresley, P.; Casper, J.; and Catanzaro, B. 2019 · 1909
Earlier work this paper cites.
Energy-Aware Neural Architecture Optimization with Fast Splitting Steepest Descent
Wang, D.; Li, M.; Wu, L.; Chandra, V.; and Liu, Q. 2019b · 1910
Earlier work this paper cites.
Wu, L.; Ye, M.; Lei, Q.; Lee, J. D.; and Liu, Q. 2020b · 2003
Earlier work this paper cites.
Heuristic Rank Selection with Progressively Searching Tensor Ring Network
Li, N.; Pan, Y.; Chen, Y.; Ding, Z.; Zhao, D.; and Xu, Z. 2020 · 2009
Earlier work this paper cites.
Progressively Stacking 2.0: A Multi-stage Layerwise Training Method for BERT Training Speedup
Yang, C.; Wang, S.; Yang, C.; Li, Y.; He, R.; and Zhang, J. 2020 · 2011
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
Zhu, Y.; Kiros, R.; Zemel, R. S.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
Net2Net: Accelerating Learning via Knowledge Transfer
Chen, T.; Goodfellow, I. J.; and Shlens, J. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Multi-level Residual Networks from Dynamical Systems View
Chang, B.; Meng, L.; Haber, E.; Tung, F.; and Begert, D. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
Universal Transformers
Dehghani, M.; Gouws, S.; Vinyals, O.; Uszkoreit, J.; and Kaiser, L. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Efficient Training of BERT by Progressively Stacking
Gong, L.; He, D.; Li, Z.; Qin, T.; Wang, L.; and Liu, T. 2019 · 2019
Cited alongside, same era.
Compressing Recurrent Neural Networks with Tensor Ring for Action Recognition
Pan, Y.; Xu, J.; Wang, M.; Ye, J.; Wang, F.; Bai, K.; and Xu, Z. 2019 · 2019
Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; and Wang, M. 2021 · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Later among the works it cites.
On the Transformer Growth for Progressive BERT Training
Gu, X.; Liu, L.; Yu, H.; Li, J.; Chen, C.; and Han, J. 2021 · 2021
Later among the works it cites.
Taking Notes on the Fly Helps Language Pre-Training
Wu, Q.; Xing, C.; Li, Y.; Ke, G.; He, D.; and Liu, T. 2021 · 2021
Later among the works it cites.
Speeding up Deep Model Training by Sharing Weights and Then Unsharing
Yang, S.; Hou, L.; Song, X.; Liu, Q.; and Zhou, D. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Cited alongside, same era.
Splitting Steepest Descent for Growing Neural Architectures
Wu, L.; Wang, D.; and Liu, Q. 2019 · 2019
Cited alongside, same era.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2020 · 2020
Cited alongside, same era.
A Primer in BERTology: What We Know About How BERT Works
Rogers, A.; Kovaleva, O.; and Rumshisky, A. 2020 · 2020
Cited alongside, same era.
Concatenated Tensor Networks for Deep Multi-Task Learning
Wang, M.; Su, Z.; Luo, X.; Pan, Y.; Zheng, S.; and Xu, Z. 2020 · 2020
Cited alongside, same era.
Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
Zhang, M.; and He, Y. 2020 · 2020
Cited alongside, same era.
Token Merging: Your ViT But Faster
Bolya, D.; Fu, C.; Dai, X.; Zhang, P.; Feichtenhofer, C.; and Hoffman, J. 2022 · 2022
Later among the works it cites.
bert2BERT: Towards Reusable Pretrained Language Models
Chen, C.; Yin, Y.; Shang, L.; Jiang, X.; Qin, Y.; Wang, F.; Wang, Z.; Chen, X.; Liu, Z.; and Liu, Q. 2022 · 2022
Later among the works it cites.
Exploring Low Rank Training of Deep Neural Networks
Kamalakara, S. R.; Locatelli, A.; Venkitesh, B.; Ba, J.; Gal, Y.; and Gomez, A. N. 2022 · 2022
Later among the works it cites.
Automated Progressive Learning for Efficient Training of Vision Transformers
Li, C.; Zhuang, B.; Wang, G.; Liang, X.; Chang, X.; and Yang, Y. 2022 · 2022
Later among the works it cites.
Knowledge Inheritance for Pre-trained Language Models
Qin, Y.; Lin, Y.; Yi, J.; Zhang, J.; Han, X.; Zhang, Z.; Su, Y.; Liu, Z.; Li, P.; Sun, M.; and Zhou, J. 2022 · 2022
Later among the works it cites.
Staged Training for Transformer Language Models
Shen, S.; Walsh, P.; Keutzer, K.; Dodge, J.; Peters, M. E.; and Beltagy, I. 2022 · 2022
Later among the works it cites.
Reusing Pretrained Models by Multi-linear Operators for Efficient Training
Pan, Y.; Yuan, Y.; Yin, Y.; Xu, Z.; Shang, L.; Jiang, X.; and Liu, Q. 2023 · 2023
Later among the works it cites.