Fetching the paper…
Reading the bibliography…
While Reinforcement Learning from Human Feedback (RLHF) has shown promise in aligning generative AI, we present empirical evidence that it can also cause severe, systematic misalignment.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Individual choice behavior , volume 4
Luce, R. D · 1959
Earlier work this paper cites.
Statistical methods for research workers
Fisher, R. A · 1970
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Hansen, E. A., Bernstein, D. S., and Zilberstein, S · 2004
Earlier work this paper cites.
Pearson’s correlation coefficient
Sedgwick, P · 2012
Earlier work this paper cites.
Corrigibility
Soares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Avoiding wireheading with value reinforcement learning
Everitt, T. and Hutter, M · 2016
Earlier work this paper cites.
Learning mixtures of plackett-luce models
Zhao, Z., Piech, P., and Xia, L · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
Everitt, T., Krakovna, V., Orseau, L., Hutter, M., and Legg, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Prolific. ac—a subject pool for online experiments
Palan, S. and Schitter, C · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
An mturk crisis? shifts in data quality and the impact on study results
Chmielewski, M. and Kucker, S. C · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Everitt, T., Hutter, M., Kumar, R., and Krakovna, V · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Bakker, M., Chadwick, M., Sheahan, H., Tessler, M., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M., et al · 2022
Earlier work this paper cites.
The expertise problem: Learning from specialized feedback
Daniels-Koch, O. and Freedman, R · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Cited alongside, same era.
On the sensitivity of reward inference to misspecified human models
Hong, J., Bhatia, K., and Dragan, A · 2022
Cited alongside, same era.
Lindner, D. and El-Assady, M · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Modeling and mitigating human annotation errors to design efficient stream processing systems with human-in-the-loop machine learning
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Openchat: Advancing open-source language models with mixed-quality data
Wang, G., Cheng, S., Zhan, X., Li, X., Song, S., and Liu, Y · 2023
Later among the works it cites.
Simple synthetic data reduces sycophancy in large language models
Wei, J., Huang, D., Lu, Y., Zhou, D., and Le, Q. V · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pandey, R., Purohit, H., Castillo, C., and Shalin, V. L · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T. T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Biyik, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D · 2023
Cited alongside, same era.
Alpagasus: Training a better alpaca with fewer data
Chen, L., Li, S., Yan, J., Wang, H., Gunaratna, K., Yadav, V., Tang, Z., Srinivasan, V., Zhou, T., Huang, H., et al · 2023
Cited alongside, same era.
Fernandes, P., Deutsch, D., Finkelstein, M., Riley, P., Martins, A. F., Neubig, G., Garg, A., Clark, J. H., Freitag, M., and Firat, O · 2023
Cited alongside, same era.
Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., et al · 2023
Later among the works it cites.
How customers are making more informed shopping decisions with rufus, amazon’s generative ai-powered shopping assistant
Amazon · 2024
Later among the works it cites.
Claude 2
Anthropic · 2024
Later among the works it cites.
Benchmarking foundation models with language-model-as-an-examiner
Bai, Y., Ying, J., Cao, Y., Lv, X., He, Y., Wang, X., Yu, J., Zeng, K., Xiao, Y., Lyu, H., et al · 2024
Later among the works it cites.
Odin: Disentangled reward mitigates hacking in rlhf
Chen, L., Zhu, C., Soselia, D., Chen, J., Zhou, T., Goldstein, T., Huang, H., Shoeybi, M., and Catanzaro, B · 2024
Later among the works it cites.
Sycophancy to subterfuge: Investigating reward-tampering in large language models
Denison, C., MacDiarmid, M., Barez, F., Duvenaud, D., Kravec, S., Marks, S., Schiefer, N., Soklaski, R., Tamkin, A., Kaplan, J., et al · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, C. X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P. S., and Hashimoto, T. B · 2024
Later among the works it cites.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al · 2024
Later among the works it cites.
Lang, L., Foote, D., Russell, S., Dragan, A., Jenner, E., and Emmons, S · 2024
Later among the works it cites.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Trustllm: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al · 2024
Later among the works it cites.
Language models learn to mislead humans via rlhf
Wen, J., Zhong, R., Khan, A., Perez, E., Steinhardt, J., Huang, M., Bowman, S. R., He, H., and Feng, S · 2024
Later among the works it cites.
On targeted manipulation and deception when optimizing llms for user feedback
Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A · 2024
Later among the works it cites.
Less: Selecting influential data for targeted instruction tuning
Xia, M., Malladi, S., Gururangan, S., Arora, S., and Chen, D · 2024
Later among the works it cites.