Open problems and fundamental limitations of reinforcement learning from human feedback
Original
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Closest in time.
Ultrafeedback: Boosting language models with high-quality feedback
Original
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M · 2023
Closest in time.
Raft: Reward ranked finetuning for generative foundation model alignment
Original
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Original
Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Closest in time.
Mistral 7b
Original
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Closest in time.
Understanding the effects of rlhf on llm generalisation and diversity
Original
Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Closest in time.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Original
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Closest in time.
Statistical rejection sampling improves preference optimization
Original
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., and Roberts, A · 2023
Closest in time.
Full parameter fine-tuning for large language models with limited resources
Original
Lv, K., Yang, Y., Liu, T., Gao, Q., Guo, Q., and Qiu, X · 2023
Closest in time.
Nash learning from human feedback
Original
Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., et al · 2023
Closest in time.
Gpt-4 technical report
Original
OpenAI · 2023
Closest in time.
Reward gaming in conditional text generation
Pang, R. Y., Padmakumar, V., Sellam, T., Parikh, A., and He, H · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Original
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2023
Closest in time.
Efficient rlhf: Reducing the memory usage of ppo
Original
Santacroce, M., Lu, Y., Yu, H., Li, Y., and Shen, Y · 2023
Closest in time.
Reward collapse in aligning large language models
Original
Song, Z., Cai, T., Lee, J. D., and Su, W. J · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Original
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Zephyr: Direct distillation of lm alignment
Original
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al · 2023
Closest in time.
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
Original
Xiong, W., Dong, H., Ye, C., Zhong, H., Jiang, N., and Zhang, T · 2023
Closest in time.
Deepspeed-chat: Easy, fast and affordable rlhf training of chatgpt-like models at all scales
Original
Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., et al · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Original
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Closest in time.
Rlhf deciphered: A critical analysis of reinforcement learning from human feedback for llms
Original
Chaudhari, S., Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A., and da Silva, B. C · 2024
Closest in time.
Break the sequential dependency of llm inference using lookahead decoding
Original
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2024
Closest in time.