Fetching the paper…
Reading the bibliography…
The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options.
English gigaword
Graff, D., Kong, J., Chen, K., and Maeda, K · 2003
Earlier work this paper cites.
Generating sequences with recurrent neural networks, 2014
Graves, A · 2014
Earlier work this paper cites.
Semi-supervised sequence learning, 2015
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Rush, A. M., Chopra, S., and Weston, J · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch, 2017
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Large scale fine-grained categorization and domain-specific transfer learning
Cui, Y., Song, Y., Sun, C., Howard, A. G., and Belongie, S. J · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Fine-tuning pre-trained transformer language models to distantly supervised relation extraction
Alt, C., Hübner, M., and Hennig, L · 2019
Earlier work this paper cites.
Plato: Pre-trained dialogue generation model with discrete latent variable
Bao, S., He, H., Wang, F., and Wu, H · 2019
Earlier work this paper cites.
Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news, 2019
Foundation, W · 2019
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
A constructive prediction of the generalization error across scales, 2019
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 2019
Earlier work this paper cites.
Transferability and hardness of supervised classification tasks
Tran, A., Nguyen, C. V., and Hassner, T · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Duality diagram similarity: a generic framework for initialization selection in task transfer learning
Dwivedi, K., Huang, J., Cichy, R. M., and Roig, G · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Leep: A new measure to evaluate transferability of learned representations
Nguyen, C. V., Hassner, T., Archambeau, C., and Seeger, M. W · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors · 2020
Cited alongside, same era.
Exploring and predicting transferability across nlp tasks
Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
Scaling laws for generative mixed-modal language models
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L · 2023
Later among the works it cites.
Bai, J., Zhang, X., Li, C., Hong, H., Xu, X., Lin, C., and Rong, W · 2023
Later among the works it cites.
Broken neural scaling laws, 2023
Caballero, E., Gupta, K., Rish, I., and Krueger, D · 2023
Later among the works it cites.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Dey, N., Gosal, G., Khachane, H., Marshall, W., Pathria, R., Tom, M., Hestness, J., et al · 2023
Later among the works it cites.
Scaling laws for multilingual neural machine translation
Fernandes, P., Ghorbani, B., Garcia, X., Freitag, M., and Firat, O · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
mt5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2020
Cited alongside, same era.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Cited alongside, same era.
A linearized framework and a new benchmark for model selection for fine-tuning
Deshpande, A., Achille, A., Ravichandran, A., Li, H., Zancato, L., Fowlkes, C., Bhotika, R., Soatto, S., and Perona, P · 2021
Cited alongside, same era.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C · 2021
Cited alongside, same era.
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Newer is not always better: Rethinking transferability metrics, their peculiarities, stability and performance
Ibrahim, S., Ponomareva, N., and Mazumder, R · 2021
Cited alongside, same era.
Transferability estimation using bhattacharyya class separability
P’andy, M., Agostinelli, A., Uijlings, J. R. R., Ferrari, V., and Mensink, T · 2021
Cited alongside, same era.
Later among the works it cites.
Scaling laws for sparsely-connected foundation models
Frantar, E., Riquelme, C., Houlsby, N., Alistarh, D., and Evci, U · 2023
Later among the works it cites.
Guided recommendation for model fine-tuning
Li, H., Fowlkes, C. C., Yang, H., Dabeer, O., Tu, Z., and Soatto, S. · 2023
Later among the works it cites.
Class incremental learning via likelihood ratio based task prediction
Lin, H., Shao, Y., Qian, W., Pan, N., Guo, Y., and Liu, B · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Later among the works it cites.
Scaling data-constrained language models
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., Kohli, P., Fawzi, A., Grochow, J., Lodi, A., Mouret, J.-B., Ringer, T., and Yu, T · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Lamini-lm: A diverse herd of distilled models from large-scale instructions
Wu, M., Waheed, A., Zhang, C., Abdul-Mageed, M., and Aji, A. F · 2023
Later among the works it cites.
Model spider: Learning to rank pre-trained models efficiently
Zhang, Y.-K., Huang, T., Ding, Y.-X., chuan Zhan, D., and Ye, H.-J · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Bi, D.-A. X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., Gao, H., Gao, K., Gao, W., Ge, R., Guan, K., Guo, D., Guo, J., Hao, G., Hao, Z., He, Y., Hu, W.-H., Huang, P., Li, E., Li, G., Li, J., Li, Y., Li, Y. K., Liang, W., Lin, F., Liu, A. X., Liu, B., Liu, W., Liu, X., Liu, X., Liu, Y., Lu, H., Lu, S., Luo, F., Ma, S., Nie, X., Pei, T., Piao, Y., Qiu, J., Qu, H., Ren, T., Ren, Z., Ruan, C., Sha, Z., Shao, Z., Song, J.-M., Su, X., Sun, J., Sun, Y., Tang, M., Wang, B.-L., Wang, P., Wang, S., Wang, Y., Wang, Y., Wu, T., Wu, Y., Xie, X., Xie, Z., Xie, Z., Xiong, Y., Xu, H., Xu, R. X., Xu, Y., Yang, D., mei You, Y., Yu, S., yuan Yu, X., Zhang, B., Zhang, H., Zhang, L., Zhang, L., Zhang, M., Zhang, M., Zhang, W., Zhang, Y., Zhao, C., Zhao, Y., Zhou, S., Zhou, S., Zhu, Q., and Zou, Y · 2024
Closest in time.
Understanding emergent abilities of language models from the loss perspective
Du, Z., Zeng, A., Dong, Y., and Tang, J · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.
Lamini-lm: A diverse herd of distilled models from large-scale instructions, 2024
Wu, M., Waheed, A., Zhang, C., Abdul-Mageed, M., and Aji, A. F · 2024
Closest in time.
When scaling meets llm finetuning: The effect of data, model and finetuning method
Zhang, B., Liu, Z., Cherry, C., and Firat, O · 2024
Closest in time.