Carbon emissions and large neural network training
Original
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., and Dean, J. (2021) · 2021
Later among the works it cites.
Combined scaling for zero-shot transfer learning
Original
Pham, H., Dai, Z., Ghiasi, G., Kawaguchi, K., Liu, H., Yu, A. W., Yu, J., Chen, Y.-T., Luong, M.-T., Wu, Y., et al. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Original
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. (2021) · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. (2021) · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Original
Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al. (2021) · 2021
Later among the works it cites.
Revisiting neural scaling laws in language and vision
Alabdulmohsin, I., Neyshabur, B., and Zhai, X. (2022) · 2022
Later among the works it cites.
Data scaling laws in NMT: The effect of noise and architecture
Original
Bansal, Y., Ghorbani, B., Garg, A., Zhang, B., Krikun, M., Cherry, C., Neyshabur, B., and Firat, O. (2022) · 2022
Later among the works it cites.
Wide attention is the way forward for transformers
Original
Brown, J. R., Zhao, Y., Shumailov, I., and Mullins, R. D. (2022) · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., Kolesnikov, A., Puigcerver, J., Ding, N., Rong, K., Akbari, H., Mishra, G., Xue, L., Thapliyal, A., Bradbury, J., Kuo, W., Seyedhosseini, M., Jia, C., Ayan, B. K., Riquelme, C., Steiner, A., Angelova, A., Zhai, X., Houlsby, N., and Soricut, R. (2022) · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Original
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. (2022) · 2022
Later among the works it cites.
The efficiency misnomer
Dehghani, M., Arnab, A., Beyer, L., Vaswani, A., and Tay, Y. (2022) · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Later among the works it cites.
UViM: A unified modeling approach for vision with learned guiding codes
Kolesnikov, A., Susano Pinto, A., Beyer, L., Zhai, X., Harmsen, J., and Houlsby, N. (2022) · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al. (2022) · 2022
Later among the works it cites.
Scaling laws from the data manifold dimension
Sharma, U. and Kaplan, J. (2022) · 2022
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A. P., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L. (2022) · 2022
Later among the works it cites.
DeiT III: Revenge of the ViT
Touvron, H., Cord, M., and Jégou, H. (2022) · 2022
Later among the works it cites.
CoCa: Contrastive captioners are image-text foundation models
Original
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y. (2022) · 2022
Later among the works it cites.
Papers With Code: ImageNet Benchmark
Code, P. W. (2023) · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Riquelme, C., Minderer, M., Puigcerver, J., Evci, U., Kumar, M., van Steenkiste, S., Elsayed, G. F., Mahendran, A., Yu, F., Oliver, A., Huot, F., Bastings, J., Collier, M. P., Gritsenko, A., Birodkar, V., Vasconcelos, C., Tay, Y., Mensink, T., Kolesnikov, A., Pavetić, F., Tran, D., Kipf, T., Lučić, M., Zhai, X., Keysers, D., Harmsen, J., and Houlsby, N. (2023) · 2023
Closest in time.
The effectiveness of mae pre-pretraining for billion-scale pretraining
Original
Singh, M., Duval, Q., Alwala, K. V., Fan, H., Aggarwal, V., Adcock, A., Joulin, A., Dollár, P., Feichtenhofer, C., Girshick, R., et al. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023) · 2023
Closest in time.