Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) aligns language models to human preferences by employing a singular reward model derived from preference data.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Aroyo, L. and Welty, C · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Collective Choice and Social Welfare
Sen, A · 2017
Earlier work this paper cites.
Towards coherent and cohesive long-form text generation
Cho, W. S., Zhang, P., Zhang, Y., Li, X., Galley, M., Brockett, C., Wang, M., and Gao, J · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
The alignment problem: Machine learning and human values
Christian, B · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2020
Earlier work this paper cites.
The state of online harassment
Vogels, E. A · 2021
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences, 2022
Bakker, M. A., Chadwick, M. J., Sheahan, H. R., Tessler, M. H., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M. M., and Summerfield, C · 2022
Cited alongside, same era.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Sap, M., Swayamdipta, S., Vianna, L., Zhou, X., Choi, Y., and Smith, N. A · 2022
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Koala: A dialogue model for academic research
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards, 2023
Ramé, A., Couairon, G., Shukor, M., Dancette, C., Gaya, J.-B., Soulier, L., and Cord, M · 2023
Later among the works it cites.
Why don’t you do it right? analysing annotators’ disagreement in subjective tasks
Sandri, M., Leonardelli, E., Tonelli, S., and Jezek, E · 2023
Later among the works it cites.
Whose opinions do language models reflect?
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D · 2023
Cited alongside, same era.
Camels in a changing climate: Enhancing lm adaptation with tulu 2
Ivison, H., Wang, Y., Pyatkin, V., Lambert, N., Peters, M., Dasigi, P., Jang, J., Wadden, D., Smith, N. A., Beltagy, I., et al · 2023
Cited alongside, same era.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Jang, J., Kim, S., Lin, B. Y., Wang, Y., Hessel, J., Zettlemoyer, L., Hajishirzi, H., Choi, Y., and Ammanabrolu, P · 2023
Cited alongside, same era.
A survey of reinforcement learning from human feedback, 2023
Kaufmann, T., Weng, P., Bengs, V., and Hüllermeier, E · 2023
Cited alongside, same era.
Large language models as superpositions of cultural perspectives, 2023
Kovač, G., Sawayama, M., Portelas, R., Colas, C., Dominey, P. F., and Oudeyer, P.-Y · 2023
Cited alongside, same era.
Reinforcement learning with human feedback: Learning dynamic choices via pessimism
Li, Z., Yang, Z., and Wang, M · 2023
Cited alongside, same era.
’generative ci’ through collective response systems, 2023
Ovadya, A · 2023
Cited alongside, same era.
The reasonable effectiveness of diverse evaluation data, 2023a
Aroyo, L., Diaz, M., Homan, C., Prabhakaran, V., Taylor, A., and Wang, D
Cited in the paper.
Later among the works it cites.
Aligning large language models with human: A survey
Wang, Y., Zhong, W., Li, L., Mi, F., Zeng, X., Huang, W., Shang, L., Jiang, X., and Liu, Q · 2023
Later among the works it cites.
Unified off-policy learning to rank: a reinforcement learning perspective
Zhang, Z., Su, Y., Yuan, H., Wu, Y., Balasubramanian, R., Wu, Q., Wang, H., and Wang, M · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons, 2023
Zhu, B., Jiao, J., and Jordan, M. I · 2023
Later among the works it cites.
Parl: A unified framework for policy alignment in reinforcement learning
Chakraborty, S., Bedi, A., Koppel, A., Wang, H., Manocha, D., Wang, M., and Huang, F · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q · 2024
Closest in time.