Fetching the paper…
Reading the bibliography…
Recently, FuseLLM introduced the concept of knowledge fusion to transfer the collective knowledge of multiple structurally varied LLMs into a target LLM through lightweight continual training.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Turc, I., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 1908
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019) · 1910
Earlier work this paper cites.
Animating rotation with quaternion curves
Shoemake, K. (1985) · 1985
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K. (1994) · 1994
Earlier work this paper cites.
Turning bayesian model averaging into bayesian model combination
Monteith, K., Carroll, J. L., Seppi, K., and Martinez, T. (2011) · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Ensemble learning: A survey
Sagi, O. and Rokach, L. (2018) · 2018
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J. (2019) · 2019
Earlier work this paper cites.
Stochastic weight averaging in parallel: Large-batch training that generalizes well
Gupta, V., Serrano, S. A., and DeCoste, D. (2020) · 2020
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q. (2020) · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M. (2020) · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2020) · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. (2021) · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Cited alongside, same era.
Mergedistill: Merging language models using pre-trained distillation
Khanuja, S., Johnson, M., and Talukdar, P. (2021) · 2021
Cited alongside, same era.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. (2022) · 2022
Cited alongside, same era.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P. (2022) · 2022
Openassistant conversations–democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., et al. (2023) · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D. (2023) · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A. (2023) · 2023
Later among the works it cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A. (2022) · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. (2022) · 2022
Cited alongside, same era.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O. (2023) · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. (2023) · 2023
Cited alongside, same era.
Specializing smaller language models towards multi-step reasoning
Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T. (2023) · 2023
Cited alongside, same era.
Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M. (2023) · 2023
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Jiang, D., Ren, X., and Lin, B. Y. (2023) · 2023
Cited alongside, same era.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
Explore-instruct: Enhancing domain-specific instruction coverage through active exploration
Wan, F., Huang, X., Yang, T., Quan, X., Bi, W., and Shi, S. (2023) · 2023
Later among the works it cites.
Openchat: Advancing open-source language models with mixed-quality data
Wang, G., Cheng, S., Zhan, X., Li, X., Song, S., and Liu, Y. (2023) · 2023
Later among the works it cites.
Magicoder: Source code is all you need
Wei, Y., Wang, Z., Liu, J., Ding, Y., and Zhang, L. (2023) · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D. (2023) · 2023
Later among the works it cites.
Ties-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M. (2023) · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023) · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024) · 2024
Closest in time.
Knowledge fusion of large language models
Wan, F., Huang, X., Cai, D., Quan, X., Bi, W., and Shi, S. (2024) · 2024
Closest in time.