Fetching the paper…
Reading the bibliography…
While training large language models (LLMs) from scratch can indeed lead to models with distinct capabilities and strengths, it incurs substantial costs and may lead to redundancy in competencies.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Turc, I., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 1908
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019) · 1910
Earlier work this paper cites.
On the mathematical foundations of theoretical statistics
Fisher, R. A. (1922) · 1922
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K. (1994) · 1994
Earlier work this paper cites.
Turning bayesian model averaging into bayesian model combination
Monteith, K., Carroll, J. L., Seppi, K., and Martinez, T. (2011) · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M. (2016) · 2016
Earlier work this paper cites.
Learning from multiple teacher networks
You, S., Xu, C., Xu, C., and Tao, D. (2017) · 2017
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2019) · 2019
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J. (2019) · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. (2020) · 2020
Earlier work this paper cites.
Stochastic weight averaging in parallel: Large-batch training that generalizes well
Gupta, V., Serrano, S. A., and DeCoste, D. (2020) · 2020
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q. (2020) · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M. (2020) · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2020) · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. (2021) · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Mergedistill: Merging language models using pre-trained distillation
Khanuja, S., Johnson, M., and Talukdar, P. (2021) · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N. (2022) · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al. (2022) · 2022
Cited alongside, same era.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A. (2022) · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. (2022) · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models
Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., and Hou, L. (2023) · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif
Zhu, B., Frick, E., Wu, T., Zhu, H., and Jiao, J. (2023) · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
Agarwal, R., Vieillard, N., Zhou, Y., Stanczyk, P., Garea, S. R., Geist, M., and Bachem, O. (2024) · 2024
Closest in time.
Evolutionary optimization of model merging recipes
Akiba, T., Shing, M., Tang, Y., Sun, Q., and Ha, D. (2024) · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al. (2023) · 2023
Cited alongside, same era.
Specializing smaller language models towards multi-step reasoning
Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T. (2023) · 2023
Cited alongside, same era.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A. (2023) · 2023
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Jiang, D., Ren, X., and Lin, B. Y. (2023) · 2023
Cited alongside, same era.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P. (2023) · 2023
Cited alongside, same era.
Solar 10.7 b: Scaling large language models with simple yet effective depth up-scaling
Kim, D., Park, C., Kim, S., Lee, W., Song, W., Kim, Y., Kim, H., Kim, Y., Lee, H., Kim, J., et al. (2023) · 2023
Cited alongside, same era.
Sparse upcycling: Training mixture-of-experts from dense checkpoints
Komatsuzaki, A., Puigcerver, J., Lee-Thorp, J., Ruiz, C. R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N. (2023) · 2023
Cited alongside, same era.
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., et al. (2024) · 2024
Closest in time.
Mastering text, code and math simultaneously via fusing highly specialized language models
Ding, N., Chen, Y., Cui, G., Lv, X., Xie, R., Zhou, B., Liu, Z., and Sun, M. (2024) · 2024
Closest in time.
Mixture-of-loras: An efficient multitask tuning for large language models
Feng, W., Hao, C., Zhang, Y., Han, Y., and Wang, H. (2024) · 2024
Closest in time.
MiniLLM: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M. (2024) · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024) · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al. (2024) · 2024
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D. (2024) · 2024
Closest in time.
Pack of llms: Model fusion at test-time via perplexity optimization
Mavromatis, C., Karypis, P., and Karypis, G. (2024) · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D. (2024) · 2024
Closest in time.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm
Sukhbaatar, S., Golovneva, O., Sharma, V., Xu, H., Lin, X. V., Rozière, B., Kahn, J., Li, D., Yih, W.-t., Weston, J., et al. (2024) · 2024
Closest in time.
Knowledge fusion of large language models
Wan, F., Huang, X., Cai, D., Quan, X., Bi, W., and Shi, S. (2024) · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y., Sun, Z., Yuan, H., Ji, K., Yang, Y., and Gu, Q. (2024) · 2024
Closest in time.
Ties-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M. (2024) · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., YU, J., Liu, Z., Zhang, Y., Kwok, J., Li, Z., Weller, A., and Liu, W. (2024) · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2024) · 2024
Closest in time.