GShard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Later among the works it cites.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Later among the works it cites.
Pay attention to mlps
Liu, H., Dai, Z., So, D., and Le, Q. V · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Original
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Later among the works it cites.
Hash layers for large sparse models
Roller, S., Sukhbaatar, S., Weston, J., et al · 2021
Later among the works it cites.
Searching for efficient transformers for language modeling
So, D., Mańke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V · 2021
Later among the works it cites.
Synthesizer: Rethinking self-attention for transformer models
Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., and Zheng, C · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Later among the works it cites.
Taming sparsely activated transformer with stochastic experts
Original
Zuo, S., Liu, X., Jiao, J., Kim, Y. J., Hassan, H., Zhang, R., Zhao, T., and Gao, J · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Original
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Original
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Transformer quality in linear time
Hua, W., Dai, Z., Liu, H., and Le, Q · 2022
Later among the works it cites.
Residual mixture of experts
Original
Wu, L., Liu, M., Chen, Y., Chen, D., Dai, X., and Yuan, L · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing, 2022
Original
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A., Chen, Z., Le, Q., and Laudon, J · 2022
Later among the works it cites.