Fetching the paper…
Reading the bibliography…
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model's fundamental understanding of specific downstream domains (e.g., math and code).
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
Principles and procedures of statistics, with special reference to the biological sciences
Carpenter, R., 1960 · 1960
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Liu, D.C., Nocedal, J., 1989 · 1989
Earlier work this paper cites.
Robust estimation of a location parameter, in: Breakthroughs in statistics: Methodology and distribution. Springer, pp. 492–518
Huber, P.J., 1992 · 1992
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., Amodei, D., 2020 · 2001
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., Smith, N.A., 2020 · 2004
Earlier work this paper cites.
Continual domain-tuning for pretrained language models
Rongali, S., Jagannatha, A., Rawat, B.P.S., Yu, H., 2020 · 2004
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M.M.A., Yang, Y., Zhou, Y., 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Earlier work this paper cites.
Ernie 2.0: A continual pre-training framework for language understanding, in: Proceedings of the AAAI conference on artificial intelligence, pp. 8968–8975
Sun, Y., Wang, S., Li, Y., Feng, S., Tian, H., Wu, H., Wang, H., 2020 · 2020
Earlier work this paper cites.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., Cherry, C., 2021 · 2021
Earlier work this paper cites.
Demix layers: Disentangling domains for modular language modeling
Gururangan, S., Lewis, M., Holtzman, A., Smith, N.A., Zettlemoyer, L., 2021 · 2021
Earlier work this paper cites.
Hernandez, D., Kaplan, J., Henighan, T., McCandlish, S., 2021 · 2021
Earlier work this paper cites.
Lifelong pretraining: Continually adapting language models to emerging corpora
Jin, X., Zhang, D., Zhu, H., Xiao, W., Li, S.W., Wei, X., Arnold, A., Ren, X., 2021 · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J.W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al., 2021 · 2021
Earlier work this paper cites.
Pretrained language model in continual learning: A comparative study, in: International conference on learning representations
Wu, T., Caccia, M., Li, Z., Li, Y.F., Qi, G., Haffari, G., 2021 · 2021
Earlier work this paper cites.
Revisiting neural scaling laws in language and vision
Alabdulmohsin, I.M., Neyshabur, B., Zhai, X., 2022 · 2022
Earlier work this paper cites.
Unified scaling laws for routed language models, in: International conference on machine learning, PMLR. pp. 4057–4086
Clark, A., de Las Casas, D., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., et al., 2022 · 2022
Earlier work this paper cites.
Continual pre-training mitigates forgetting in language and vision
Cossu, A., Tuytelaars, T., Carta, A., Passaro, L., Lomonaco, V., Bacciu, D., 2022 · 2022
Earlier work this paper cites.
Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset
Henderson, P., Krass, M., Zheng, L., Guha, N., Manning, C.D., Jurafsky, D., Ho, D., 2022 · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D.d.L., Hendricks, L.A., Welbl, J., Clark, A., et al., 2022 · 2022
Cited alongside, same era.
Scaling laws for generative mixed-modal language models, in: International Conference on Machine Learning, PMLR. pp. 265–279
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., Zettlemoyer, L., 2023 · 2023
Cited alongside, same era.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M.D., McAleer, S., Jiang, A.Q., Deng, J., Biderman, S., Welleck, S., 2023 · 2023
Cited alongside, same era.
Thakur, V., 2023 · 2023
Later among the works it cites.
Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models
Wang, Z.M., Peng, Z., Que, H., Liu, J., Zhou, W., Wu, Y., Guo, H., Gan, R., Ni, Z., Zhang, M., Zhang, Z., Ouyang, W., Xu, K., Chen, W., Fu, J., Peng, J., 2023 · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., Yang, F., Deng, F., Wang, F., Liu, F., Ai, G., Dong, G., Zhao, H., Xu, H., Sun, H., Zhang, H., Liu, H., Ji, J., Xie, J., Dai, J., Fang, K., Su, L., Song, L., Liu, L., Ru, L., Ma, L., Wang, M., Liu, M., Lin, M., Nie, N., Guo, P., Sun, R., Zhang, T., Li, T., Li, T., Cheng, W., Chen, W., Zeng, X., Wang, X., Chen, X., Men, X., Yu, X., Pan, X., Shen, Y., Wang, Y., Li, Y., Jiang, Y., Gao, Y., Zhang, Y., Zhou, Z., Wu, Z., 2023 · 2023
Later among the works it cites.
Analyzing the impact of data selection and fine-tuning on economic and political biases in llms
Agiza, A., Mostagir, M., Reda, S., 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al., 2023 · 2023
Cited alongside, same era.
Lifelong language pretraining with distribution-specialized experts, in: International Conference on Machine Learning, PMLR. pp. 5383–5395
Chen, W., Zhou, Y., Du, N., Huang, Y., Laudon, J., Chen, Z., Cui, C., 2023 · 2023
Cited alongside, same era.
Gpts are gpts: An early look at the labor market impact potential of large language models
Eloundou, T., Manning, S., Mishkin, P., Rock, D., 2023 · 2023
Cited alongside, same era.
Scaling laws for sparsely-connected foundation models
Frantar, E., Riquelme, C., Houlsby, N., Alistarh, D., Evci, U., 2023 · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization, in: International Conference on Machine Learning, PMLR. pp. 10835–10866
Gao, L., Schulman, J., Hilton, J., 2023 · 2023
Cited alongside, same era.
Owl: A large language model for it operations
Guo, H., Yang, J., Liu, J., Yang, L., Chai, L., Bai, J., Peng, J., Hu, X., Chen, C., Zhang, D., et al., 2023 · 2023
Cited alongside, same era.
Continual pre-training of large language models: How to (re) warm your model?
Gupta, K., Thérien, B., Ibrahim, A., Richter, M.L., Anthony, Q., Belilovsky, E., Rish, I., Lesort, T., 2023 · 2023
Cited alongside, same era.
Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., Casas, D.d.l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al., 2023 · 2023
Cited alongside, same era.
Closest in time.
Yi: Open foundation models by 01.ai
AI, ., :, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., Dai, Z., 2024 · 2024
Closest in time.
Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
Bai, G., Liu, J., Bu, X., He, Y., Liu, J., Zhou, Z., Lin, Z., Su, W., Ge, T., Zheng, B., Ouyang, W., 2024 · 2024
Closest in time.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al., 2024 · 2024
Closest in time.
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., Dong, X., Duan, H., Fan, Q., Fei, Z., Gao, Y., Ge, J., Gu, C., Gu, Y., Gui, T., Guo, A., Guo, Q., He, C., Hu, Y., Huang, T., Jiang, T., Jiao, P., Jin, Z., Lei, Z., Li, J., Li, J., Li, L., Li, S., Li, W., Li, Y., Liu, H., Liu, J., Hong, J., Liu, K., Liu, K., Liu, X., Lv, C., Lv, H., Lv, K., Ma, L., Ma, R., Ma, Z., Ning, W., Ouyang, L., Qiu, J., Qu, Y., Shang, F., Shao, Y., Song, D., Song, Z., Sui, Z., Sun, P., Sun, Y., Tang, H., Wang, B., Wang, G., Wang, J., Wang, J., Wang, R., Wang, Y., Wang, Z., Wei, X., Weng, Q., Wu, F., Xiong, Y., Xu, C., Xu, R., Yan, H., Yan, Y., Yang, X., Ye, H., Ying, H., Yu, J., Yu, J., Zang, Y., Zhang, C., Zhang, L., Zhang, P., Zhang, P., Zhang, R., Zhang, S., Zhang, S., Zhang, W., Zhang, W., Zhang, X., Zhang, X., Zhao, H., Zhao, Q., Zhao, X., Zhou, F., Zhou, Z., Zhuo, J., Zou, Y., Qiu, X., Qiao, Y., Lin, D., 2024 · 2024
Closest in time.
Chinese tiny llm: Pretraining a chinese-centric large language model
Du, X., Yu, Z., Gao, S., Pan, D., Cheng, Y., Ma, Z., Yuan, R., Qu, X., Liu, J., Zheng, T., Luo, X., Zhou, G., Yuan, B., Chen, W., Fu, J., Zhang, G., 2024 · 2024
Closest in time.
E2-llm: Efficient and extreme length extension of large language models
Liu, J., Bai, Z., Zhang, Y., Zhang, C., Zhang, Y., Zhang, G., Wang, J., Que, H., Chen, Y., Su, W., et al., 2024 · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation
Lozhkov, A., Li, R., Allal, L.B., Cassano, F., Lamy-Poirier, J., Tazi, N., Tang, A., Pykhtar, D., Liu, J., Wei, Y., et al., 2024 · 2024
Closest in time.
More human than human: Measuring chatgpt political bias
Motoki, F., Pinho Neto, V., Rodrigues, V., 2024 · 2024
Closest in time.
Scaling data-constrained language models
Muennighoff, N., Rush, A., Barak, B., Le Scao, T., Tazi, N., Piktus, A., Pyysalo, S., Wolf, T., Raffel, C.A., 2024 · 2024
Closest in time.
Mupt: A generative symbolic music pretrained transformer
Qu, X., Bai, Y., Ma, Y., Zhou, Z., Lo, K.M., Liu, J., Yuan, R., Min, L., Liu, X., Zhang, T., et al., 2024 · 2024
Closest in time.
Unicoder: Scaling code large language model via universal code
Sun, T., Chai, L., Jian Yang, Y.Y., Guo, H., Liu, J., Wang, B., Yang, L., Li, Z., 2024 · 2024
Closest in time.
Conceptmath: A bilingual concept-wise benchmark for measuring mathematical reasoning of large language models
Wu, Y., Liu, J., Bu, X., Liu, J., Zhou, Z., Zhang, Y., Zhang, C., Bai, Z., Chen, H., Ge, T., Ouyang, W., Su, W., Zheng, B., 2024 · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, Z., Zhang, Y., 2024 · 2024
Closest in time.
Data mixing laws: Optimizing data mixtures by predicting language modeling performance
Ye, J., Liu, P., Sun, T., Zhou, Y., Zhan, J., Qiu, X., 2024 · 2024
Closest in time.