Improving transformer optimization through better initialization
Xiao Shi Huang, Felipe Perez, Jimmy Ba, and Maksims Volkovs · 2020
Later among the works it cites.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Adapterfusion: Non-destructive task composition for transfer learning
Original
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
Original
Hao Tan and Mohit Bansal · 2020
Later among the works it cites.
Progressively stacking 2.0: A multi-stage layerwise training method for bert training speedup
Original
Cheng Yang, Shengnan Wang, Chao Yang, Yuechuan Li, Ru He, and Jingqiao Zhang · 2020
Later among the works it cites.
Accelerating training of transformer-based language models with progressive layer dropping
Minjia Zhang and Yuxiong He · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan · 2021
Later among the works it cites.
bert2bert: Towards reusable pretrained language models
Original
Cheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, and Qun Liu · 2021
Later among the works it cites.
Low-rank constraints for fast inference in structured models
Justin Chiu, Yuntian Deng, and Alexander Rush · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Original
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Knowledge inheritance for pre-trained language models
Original
Yujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, et al · 2021
Later among the works it cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus · 2021
Later among the works it cites.
Firefly neural architecture descent: a general approach for growing neural networks
Lemeng Wu, Dilin Wang, Peter Stone, and Qiang Liu · 2021
Later among the works it cites.
Gradinit: Learning to initialize neural networks for stable and efficient training
Chen Zhu, Renkun Ni, Zheng Xu, Kezhi Kong, W Ronny Huang, and Tom Goldstein · 2021
Later among the works it cites.
Monarch: Expressive structured matrices for efficient and accurate training
Tri Dao, Beidi Chen, Nimit S Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher Ré · 2022
Later among the works it cites.
Gradmax: Growing neural networks using gradient information
Original
Utku Evci, Max Vladymyrov, Thomas Unterthiner, Bart van Merriënboer, and Fabian Pedregosa · 2022
Later among the works it cites.
Token dropping for efficient bert pretraining
Original
Le Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu, Xinying Song, Xiaodan Song, and Denny Zhou · 2022
Later among the works it cites.
Visual prompt tuning
Original
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim · 2022
Later among the works it cites.
Automated progressive learning for efficient training of vision transformers
Changlin Li, Bohan Zhuang, Guangrun Wang, Xiaodan Liang, Xiaojun Chang, and Yi Yang · 2022
Later among the works it cites.
Staged training for transformer language models
Original
Sheng Shen, Pete Walsh, Kurt Keutzer, Jesse Dodge, Matthew Peters, and Iz Beltagy · 2022
Later among the works it cites.