How many layers and why? An analysis of the model depth in transformers
A. Simoulin and B. Crabbé · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
W. Fedus, B. Zoph, and N. Shazeer · 2022
Later among the works it cites.
Longt5: Efficient text-to-text transformer for long sequences, 2022
M. Guo, J. Ainslie, D. Uthus, S. Ontanon, J. Ni, Y.-H. Sung, and Y. Yang · 2022
Later among the works it cites.
Confident adaptive language modeling, 2022
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Q. Tran, Y. Tay, and D. Metzler · 2022
Later among the works it cites.
St-moe: Designing stable and transferable sparse expert models, 2022
B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus · 2022
Later among the works it cites.
Colt5: Faster long-range transformers with conditional computation, 2023
J. Ainslie, T. Lei, M. de Jong, S. Ontañón, S. Brahma, Y. Zemlyanskiy, D. Uthus, M. Guo, J. Lee-Thorp, Y. Tay, Y.-H. Sung, and S. Sanghai · 2023
Later among the works it cites.
Token merging: Your vit but faster, 2023
D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman · 2023
Later among the works it cites.
Conditional adapters: Parameter-efficient transfer learning with fast inference, 2023
T. Lei, J. Bai, S. Brahma, J. Ainslie, K. Lee, Y. Zhou, N. Du, V. Y. Zhao, Y. Wu, B. Li, Y. Zhang, and M.-W. Chang · 2023
Later among the works it cites.