Fetching the paper…
Reading the bibliography…
A widespread strategy to obtain a language model that performs well on a target domain is to finetune a pretrained model to perform unsupervised next-token prediction on data from that target domain.
Parameter-efficient transfer learning for nlp, 2019
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 1902
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I · 2014
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction, 2017
Hastie, T., Tibshirani, R., and Friedman, J · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Kudo, T · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Scaling laws for transfer, 2021
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Self-attention does not need o ( n 2 ) o(n2) memory
Rabe, M. N. and Staats, C · 2021
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning, 2022
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2022
Cited alongside, same era.
Improved fine-tuning by better leveraging pre-training data
Liu, Z., Xu, Y., Xu, Y., Qian, Q., Li, H., Ji, X., Chan, A., and Jin, R · 2022
Cited alongside, same era.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., et al · 2022
Cited alongside, same era.
Disentangling and mitigating the impact of task similarity for continual learning, 2024
Hiratani, N · 2024
Later among the works it cites.
MiniCPM: Unveiling the potential of small language models with scalable training strategies
Hu, S., Tu, Y., Han, X., Cui, G., He, C., Zhao, W., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhang, X., Thai, Z. L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M · 2024
Later among the works it cites.
Simple and scalable strategies to continually pre-train large language models, 2024
Ibrahim, A., Thérien, B., Gupta, K., Richter, M. L., Anthony, Q., Lesort, T., Belilovsky, E., and Rish, I · 2024
Later among the works it cites.
Scaling laws for downstream task performance of large language models
Isik, B., Ponomareva, N., Hazimeh, H., Paparas, D., Vassilvitskii, S., and Koyejo, S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
Scaling laws for generative mixed-modal language models
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L · 2023
Cited alongside, same era.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2023
Cited alongside, same era.
Scaling data-constrained language models
Muennighoff, N., Rush, A., Barak, B., Le Scao, T., Tazi, N., Piktus, A., Pyysalo, S., Wolf, T., and Raffel, C. A · 2023
Cited alongside, same era.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Teknium · 2023
Cited alongside, same era.
Physics in next-token prediction, 2024
An, H., Song, Y., and Li, X · 2024
Cited alongside, same era.
An empirical study of scaling laws for transfer
Barnett, M · 2024
Cited alongside, same era.
Kalajdzievski, D · 2024
Later among the works it cites.
Get more for less: Principled data selection for warming up fine-tuning in llms
Kang, F., Just, H. A., Sun, Y., Jahagirdar, H., Zhang, Y., Du, R., Sahu, A. K., and Jia, R · 2024
Later among the works it cites.
Resolving discrepancies in compute-optimal scaling of language models, 2024
Porian, T., Wortsman, M., Jitsev, J., Schmidt, L., and Carmon, Y · 2024
Later among the works it cites.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws, 2024
Sardana, N., Portes, J., Doubov, S., and Frankle, J · 2024
Later among the works it cites.
Scaling law with learning rate annealing
Tissue, H., Wang, V., and Wang, L · 2024
Later among the works it cites.
When precision meets position: Bfloat16 breaks down rope in long-context training
Wang, H., Liu, Q., Du, C., Zhu, T., Du, C., Kawaguchi, K., and Pang, T · 2024
Later among the works it cites.
Redpajama: an open dataset for training large language models
Weber, M., Fu, D. Y., Anthony, Q., Oren, Y., Adams, S., Alexandrov, A., Lyu, X., Nguyen, H., Yao, X., Adams, V., Athiwaratkun, B., Chalamala, R., Chen, K., Ryabinin, M., Dao, T., Liang, P., Ré, C., Rish, I., and Zhang, C · 2024
Later among the works it cites.
What makes a high-quality training dataset for large language models: A practitioners’ perspective
Yu, X., Zhang, Z., Niu, F., Hu, X., Xia, X., and Grundy, J · 2024
Later among the works it cites.
When scaling meets LLM finetuning: The effect of data, model and finetuning method
Zhang, B., Liu, Z., Cherry, C., and Firat, O · 2024
Later among the works it cites.
Asymmetry in low-rank adapters of foundation models, 2024
Zhu, J., Greenewald, K., Nadjahi, K., de Ocáriz Borde, H. S., Gabrielsson, R. B., Choshen, L., Ghassemi, M., Yurochkin, M., and Solomon, J · 2024
Later among the works it cites.
Llms on the line: Data determines loss-to-loss scaling laws
Mayilvahanan, P., Wiedemer, T., Mallick, S., Bethge, M., and Brendel, W · 2025
Closest in time.