Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Later among the works it cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Later among the works it cites.
Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2021
Later among the works it cites.
Pangu-alpha: Large-scale autoregressive pretrained chinese language models with auto-parallel computation
Original
Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., Jiang, X., Yang, Z., Wang, K., Zhang, X., et al · 2021
Later among the works it cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2021
Later among the works it cites.
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
Original
Abadji, J., Ortiz Suarez, P., Romary, L., and Sagot, B · 2022
Later among the works it cites.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al · 2022
Later among the works it cites.
Quantifying memorization across neural language models
Original
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Original
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Re, C · 2022
Later among the works it cites.
Scaling laws and interpretability of learning from repeated data
Original
Hernandez, D., Brown, T., Conerly, T., DasSarma, N., Drain, D., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Henighan, T., Hume, T., et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Original
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Kreutzer, J., Caswell, I., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A. A., Subramani, N., Sokolov, A., Sikasote, C., et al · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Laurençon, H., Saulnier, L., Wang, T., Akiki, C., del Moral, A. V., Le Scao, T., Von Werra, L., Mou, C., Ponferrada, E. G., Nguyen, H., et al · 2022
Later among the works it cites.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2022
Later among the works it cites.
Compute trends across three eras of machine learning
Original
Sevilla, J., Heim, L., Ho, A., Besiroglu, T., Hobbhahn, M., and Villalobos, P · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Original
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Original
Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., and Ho, A · 2022
Later among the works it cites.
What language model architecture and pretraining objective work best for zero-shot generalization?
Wang, T., Roberts, A., Hesslow, D., Scao, T. L., Chung, H. W., Beltagy, I., Launay, J., and Raffel, C · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Original
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A. K. M., Tirumala, K., Simig, D., Ganguli, S., and Morcos, A. S · 2023
Closest in time.
Luminous: performance benchmarks
Original
Aleph Alpha · 2023
Closest in time.
Falcon-40b: an open large language model with state-of-the-art performance
Almazrouei, E., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Original
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Closest in time.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Original
Dey, N., Gosal, G., Khachane, H., Marshall, W., Pathria, R., Tom, M., Hestness, J., et al · 2023
Closest in time.
Ethnologue: Languages of the World
Eberhard, D. M., Simons, G. F., and Fennig, C. D · 2023
Closest in time.
Llama: Open and efficient foundation language models
Original
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.