Fetching the paper…
Reading the bibliography…
The success of reinforcement learning from human feedback (RLHF) in language model alignment is strongly dependent on the quality of the underlying reward model.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Probability of error of some adaptive pattern-recognition machines
Scudder, H · 1965
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
Yarowsky, D · 1995
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Revisiting self-training for neural sequence generation
He, J., Gu, J., Shen, J., and Ranzato, M · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Earlier work this paper cites.
Rethinking pre-training and self-training
Zoph, B., Ghiasi, G., Lin, T.-Y., Cui, Y., Liu, H., Cubuk, E. D., and Le, Q · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Amini, M.-R., Feofanov, V., Pauletto, L., Devijver, E., and Maximov, Y · 2022
Cited alongside, same era.
Unsupervised prompt learning for vision-language models
Huang, T., Chu, J., and Wei, F · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback, 2023
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment, 2023
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T · 2023
Later among the works it cites.
Some things are more cringe than others: Preference optimization with the pairwise cringe loss, 2023
Xu, J., Lee, A., Sukhbaatar, S., and Weston, J · 2023
Later among the works it cites.
Rlcd: Reinforcement learning from contrast distillation for language model alignment
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Google · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Cited alongside, same era.
Aligning large language models through synthetic feedback
Kim, S., Bae, S., Shin, J., Kang, S., Kwak, D., Yoo, K. M., and Seo, M · 2023
Cited alongside, same era.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al
Cited in the paper.
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
Self-rewarding language models, 2024
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.