Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) aims to align language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to fine-tune the base LMs.
Fine-tuning language models from human preferences
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; and Irving, G. 2019 · 1909
Earlier work this paper cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swayamdipta, S.; Schwartz, R.; Lourie, N.; Wang, Y.; Hajishirzi, H.; Smith, N. A.; and Choi, Y. 2020 · 2009
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C.; Akrour, R.; Neumann, G.; and Fürnkranz, J. 2017 · 2017
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A.; Bai, Y.; Chen, A.; Drain, D.; Ganguli, D.; Henighan, T.; Jones, A.; Joseph, N.; Mann, B.; DasSarma, N.; et al. 2021 · 2021
Earlier work this paper cites.
He, P.; Gao, J.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. 2022 · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A.; McAleese, N.; Trebacz, M.; Aslanides, J.; Firoiu, V.; Ewalds, T.; Rauh, M.; Weidinger, L.; Chadwick, M.; Thacker, P.; et al. 2022 · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
synthetic-instruct-gptj-pairwise (Revision cc92d8d)
Alex Havrilla. 2023 · 2023
Cited alongside, same era.
Peering through preferences: Unraveling feedback acquisition for aligning large language models
Bansal, H.; Dang, J.; and Grover, A. 2023 · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S.; Davies, X.; Shi, C.; Gilbert, T. K.; Scheurer, J.; Rando, J.; Freedman, R.; Korbak, T.; Lindner, D.; Freire, P.; et al. 2023 · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L.; Schulman, J.; and Hilton, J. 2023 · 2023
Cited alongside, same era.
High-Dimension Human Value Representation in Large Language Models
Cahyawijaya, S.; Chen, D.; Bang, Y.; Khalatbari, L.; Wilie, B.; Ji, Z.; Ishii, E.; and Fung, P. 2024 · 2024
Closest in time.
Social choice for AI alignment: Dealing with diverse human feedback
Conitzer, V.; Freedman, R.; Heitzig, J.; Holliday, W. H.; Jacobs, B. M.; Lambert, N.; Mossé, M.; Pacuit, E.; Russell, S.; Schoelkopf, H.; et al. 2024 · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y.; Li, C. X.; Taori, R.; Zhang, T.; Gulrajani, I.; Ba, J.; Guestrin, C.; Liang, P. S.; and Hashimoto, T. B. 2024 · 2024
Closest in time.
Inverse Constitutional AI: Compressing Preferences into Principles
Findeis, A.; Kaufmann, T.; Hüllermeier, E.; Albanie, S.; and Mullins, R. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ChatGPT outperforms crowd workers for text-annotation tasks
Gilardi, F.; Alizadeh, M.; and Kubli, M. 2023 · 2023
Cited alongside, same era.
Reward reports for reinforcement learning
Gilbert, T. K.; Lambert, N.; Dean, S.; Zick, T.; Snoswell, A.; and Mehta, S. 2023 · 2023
Cited alongside, same era.
Human feedback is not gold standard
Hosking, T.; Blunsom, P.; and Bartolo, M. 2023 · 2023
Cited alongside, same era.
The alignment ceiling: Objective mismatch in reinforcement learning from human feedback
Lambert, N.; and Calandra, R. 2023 · 2023
Cited alongside, same era.
The history and risks of reinforcement learning and human feedback
Lambert, N.; Krendl Gilbert, T.; and Zick, T. 2023 · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X.; Zhang, T.; Dubois, Y.; Taori, R.; Gulrajani, I.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Cited alongside, same era.
Style over substance: Evaluation biases for large language models
Wu, M.; and Aji, A. F. 2023 · 2023
Cited alongside, same era.
Ge, L.; Halpern, D.; Micha, E.; Procaccia, A. D.; Shapira, I.; Vorobeychik, Y.; and Wu, J. 2024 · 2024
Closest in time.
Kirk, H. R.; Whitefield, A.; Röttger, P.; Bean, A.; Margatina, K.; Ciro, J.; Mosquera, R.; Bartolo, M.; Williams, A.; He, H.; et al. 2024 · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Chandu, K.; Dziri, N.; Kumar, S.; Zick, T.; Choi, Y.; et al. 2024 · 2024
Closest in time.
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
Multi-Granularity, M.-L. M.-F. 2024 · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Niu, Z.; Ren, H.; Gao, X.; Hua, G.; and Jin, R. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling
Wang, B.; Zheng, R.; Chen, L.; Liu, Y.; Dou, S.; Huang, C.; Shen, W.; Jin, S.; Zhou, E.; Shi, C.; et al. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Closest in time.