Fetching the paper…
Reading the bibliography…
This work identifies the Energy Loss Phenomenon in Reinforcement Learning from Human Feedback (RLHF) and its connection to reward hacking.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2009
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W.-t., and Zweig, G · 2013
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Position bias estimation for unbiased learning to rank in personal search
Wang, X., Golbandi, N., Bendersky, M., Metzler, D., and Najork, M · 2018
Earlier work this paper cites.
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D · 2019
Earlier work this paper cites.
On variational bounds of mutual information
Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G · 2019
Earlier work this paper cites.
Bert has a moral compass: Improvements of ethical and moral values of machines
Schramowski, P., Turan, C., Jentzsch, S., Rothkopf, C., and Kersting, K · 2019
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
The internal state of an llm knows when it’s lying
Azaria, A. and Mitchell, T · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al · 2023
Cited alongside, same era.
Chen, Y., Wang, R., Jiang, H., Shi, S., and Xu, R · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Reward model ensembles help mitigate overoptimization
Coste, T., Anwar, U., Kirk, R., and Krueger, D · 2024
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, C. X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P. S., and Hashimoto, T. B · 2024
Later among the works it cites.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A. N., Dvijotham, K. D., Fisch, A., Heller, K. A., Pfohl, S. R., Ramachandran, D., Shaw, P., and Berant, J · 2024
Later among the works it cites.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Later among the works it cites.
InfoRM: Mitigating reward hacking in RLHF via information-theoretic reward modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Shen, W., Zheng, R., Zhan, W., Zhao, J., Dou, S., Gui, T., Zhang, Q., and Huang, X · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles
Zhai, Y., Zhang, H., Lei, Y., Yu, Y., Xu, K., Feng, D., Ding, B., and Wang, H · 2023
Cited alongside, same era.
Delve into PPO: Implementation matters for stable RLHF
Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Zhou, Y., Xiong, L., Chen, L., Xi, Z., Xu, N., Lai, W., Zhu, M., Huang, H., Gui, T., Zhang, Q., and Huang, X · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Cited alongside, same era.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al · 2024
Cited alongside, same era.
Miao, Y., Zhang, S., Ding, L., Bao, R., Zhang, L., and Tao, D · 2024
Later among the works it cites.
Confronting reward model overoptimization with constrained RLHF
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A., and McAleer, S. M · 2024
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V · 2024
Later among the works it cites.
Scaling laws for reward model overoptimization in direct alignment algorithms
Rafailov, R., Chittepu, Y., Park, R., Sikchi, H., Hejna, J., Knox, W. B., Finn, C., and Niekum, S · 2024
Later among the works it cites.
WARM: On the benefits of weight averaged reward models
Rame, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J · 2024
Later among the works it cites.
A long way to go: Investigating length correlations in RLHF
Singhal, P., Goyal, T., Xu, J., and Durrett, G · 2024
Later among the works it cites.
Secrets of rlhf in large language models part ii: Reward modeling
Wang, B., Zheng, R., Chen, L., Liu, Y., Dou, S., Huang, C., Shen, W., Jin, S., Zhou, E., Shi, C., et al · 2024
Later among the works it cites.
Improving generalization of alignment with human preferences through group invariant learning
Zheng, R., Shen, W., Hua, Y., Lai, W., Dou, S., Zhou, Y., Xi, Z., Wang, X., Huang, H., Gui, T., et al · 2024
Later among the works it cites.
Intention analysis makes llms a good jailbreak defender
Zhang, Y., Ding, L., Zhang, L., and Tao, D · 2025
Closest in time.