Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Later among the works it cites.
Optimising human-ai collaboration by learning convincing explanations
Chan, A., Hüyük, A., and van der Schaar, M · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
Reward model ensembles help mitigate overoptimization
Original
Coste, T., Anwar, U., Kirk, R., and Krueger, D · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Original
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Original
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Later among the works it cites.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini-Team, G. D · 2023
Later among the works it cites.
Openllama: An open reproduction of llama, 2023
Geng, X. and Liu, H · 2023
Later among the works it cites.
Koala: A dialogue model for academic research
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D · 2023
Later among the works it cites.
Neural laplace control for continuous-time delayed systems
Holt, S., Hüyük, A., Qian, Z., Sun, H., and van der Schaar, M · 2023
Later among the works it cites.
Openassistant conversations–democratizing large language model alignment
Original
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., et al · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Original
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Later among the works it cites.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Let’s verify step by step
Original
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Original
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Stabilizing rlhf through advantage model and selective rehearsal
Original
Peng, B., Song, L., Tian, Y., Jin, L., Mi, H., and Yu, D · 2023
Later among the works it cites.
A survey of temporal credit assignment in deep reinforcement learning
Original
Pignatelli, E., Ferret, J., Geist, M., Mesnard, T., van Hasselt, H., and Toni, L · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Original
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Vanishing gradients in reinforcement finetuning of language models
Original
Razin, N., Zhou, H., Saremi, O., Thilak, V., Bradley, A., Nakkiran, P., Susskind, J., and Littwin, E · 2023
Later among the works it cites.
Redpajama-data: An open source recipe to reproduce llama training dataset, 2023
TogetherComputer · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Original
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Original
Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Original
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Original
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Later among the works it cites.
Secrets of rlhf in large language models part i: Ppo
Original
Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., et al · 2023
Later among the works it cites.