Cm3: A causal masked multimodal model of the internet
Original
Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., et al · 2022
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Later among the works it cites.
Data distributional properties drive emergent few-shot learning in transformers
Chan, S. C., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A., Richemond, P. H., McClelland, J., and Hill, F · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Original
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Original
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Original
Ding, M., Zheng, W., Hong, W., and Tang, J · 2022
Later among the works it cites.
Magma–multimodal augmentation of generative models through adapter-based finetuning
Eichenberg, C., Black, S., Weinbach, S., Parcalabescu, L., and Frank, A · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Pretrained transformers as universal computation engines
Lu, K., Grover, A., Abbeel, P., and Mordatch, I · 2022
Later among the works it cites.
Linearly mapping from image to text space
Original
Merullo, J., Castricato, L., Eickhoff, C., and Pavlick, E · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Original
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model
Original
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al · 2022
Later among the works it cites.
Transcending scaling laws with 0.1% extra compute
Original
Tay, Y., Wei, J., Chung, H. W., Tran, V. Q., So, D. R., Shakeri, S., Garcia, X., Zheng, H. S., Rao, J., Chowdhery, A., et al · 2022
Later among the works it cites.
Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Later among the works it cites.
Re3: Generating longer stories with recursive reprompting and revision
Yang, K., Peng, N., Tian, Y., and Klein, D · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Original
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.