Fetching the paper…
Reading the bibliography…
While protein language models (pLMs) have transformed biological research, the scaling laws governing their improvement remain underexplored.
Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Suzek, B. E., Wang, Y., Huang, H., McGarvey, P. B., Wu, C. H., and Consortium, U · 2015
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets
Steinegger, M. and Söding, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., et al · 2021
Cited alongside, same era.
Proteinbert: a universal deep-learning model of protein sequence and function
Brandes, N., Ofer, D., Peleg, Y., Rappoport, N., and Linial, M · 2022
Cited alongside, same era.
Rita: a study on scaling up generative protein sequence models
Hesslow, D., Zanichelli, N., Notin, P., Poli, I., and Marks, D · 2022
Cited alongside, same era.
Ankh: Optimized protein language model unlocks general-purpose modelling
Elnaggar, A., Essam, H., Salah-Eldin, W., Moustafa, W., Elkerdawy, M., Rochereau, C., and Rost, B · 2023
Later among the works it cites.
Evolutionary-scale prediction of atomic-level protein structure with a language model
Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., et al · 2023
Later among the works it cites.
xtrimopglm: unified 100b-scale pre-trained transformer for deciphering the language of protein
Chen, B., Cheng, X., Li, P., Geng, Y.-a., Gong, J., Li, S., Bei, Z., Tan, X., Wang, B., Zeng, X., et al · 2024
Closest in time.
Training compute-optimal protein language models
Cheng, X., Chen, B., Li, P., Gong, J., Tang, J., and Song, L · 2024
Closest in time.
Cramming protein language model training in 24 gpu hours
Frey, N. C., Joren, T., Ismail, A., Goodman, A., Bonneau, R., Cho, K., and Gligorijevic, V · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Closest in time.