Fetching the paper…
Reading the bibliography…
Heterogeneous model fusion enhances the performance of LLMs by integrating the knowledge and capabilities of multiple structurally diverse models.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Ensemble approach for natural language question answering problem
Aniol, A., Pietron, M., and Duda, J · 2019
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Kool, W., van Hoof, H., and Welling, M · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
TRL: Transformer reinforcement learning, 2020
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., and Gallouédec, Q · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Earlier work this paper cites.
Achiam, O. J., Adler, S., and Sandhini Agarwal, e. a · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
AlpacaEval: An automatic evaluator of instruction-following models, 2023
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Cited alongside, same era.
Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
ORPO: Monolithic preference optimization without reference model
Hong, J., Lee, N., and Thorne, J · 2024
Later among the works it cites.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S. I., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Later among the works it cites.
RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback
Lee, H., Phatale, S., Mansoor, H., Mesnard, T., Ferret, J., Lu, K. R., Bishop, C., Hall, E., Carbune, V., Rastogi, A., and Prakash, S · 2024
Later among the works it cites.
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., and Stoica, I · 2024
Later among the works it cites.
Statistical rejection sampling improves preference optimization
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evolutionary optimization of model merging recipes
Akiba, T., Shing, M., Tang, Y., Sun, Q., and Ha, D · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J. E., et al · 2024
Cited alongside, same era.
UltraFeedback: Boosting language models with high-quality feedback
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Mastering text, code and math simultaneously via fusing highly specialized language models
Ding, N., Chen, Y., Cui, G., Lv, X., Xie, R., Zhou, B., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., and Abhinav Pandey, e. a · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Dubois, Y., Liang, P., and Hashimoto, T · 2024
Cited alongside, same era.
KTO: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Cited alongside, same era.
Later among the works it cites.
SimPO: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Later among the works it cites.
Gpt-4o system card, 2024
OpenAI · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Riviere, G. T. M., Pathak, S., and Pier Giuseppe Sessa, e. a · 2024
Later among the works it cites.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T · 2024
Later among the works it cites.
DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model
Shao, Z., Dai, D., Guo, D., Liu), B. L. B., Wang, Z., and Xin, H · 2024
Later among the works it cites.
MuSR: Testing the limits of chain-of-thought with multistep soft reasoning
Sprague, Z. R., Ye, X., Bostrom, K., Chaudhuri, S., and Durrett, G · 2024
Later among the works it cites.
Branch-Train-MiX: Mixing expert LLMs into a mixture-of-experts LLM
Sukhbaatar, S., Golovneva, O., Sharma, V., Xu, H., Lin, X. V., Roziere, B., Kahn, J., Li, S.-W., tau Yih, W., Weston, J. E., and Li, X · 2024
Later among the works it cites.
Bridging the gap between different vocabularies for LLM ensemble
Xu, Y., Lu, J., and Zhang, J · 2024
Later among the works it cites.