Fetching the paper…
Reading the bibliography…
When applying reinforcement learning from human feedback (RLHF), the reward is learned from data and, therefore, always has some error.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019) · 1907
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2020) · 1909
Earlier work this paper cites.
‘improving ratings’: audit in the british university system
Strathern, M. (1997) · 1997
Earlier work this paper cites.
An Introduction to Heavy-Tailed and Subexponential Distributions
Foss, S., Korshunov, D., and Zachary, S. (2013) · 2013
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M. (2018) · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2018) · 2018
Cited alongside, same era.
Zhuang, S. and Hadfield-Menell, D. (2021) · 2021
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J. (2023) · 2023
Cited alongside, same era.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H. (2023) · 2023
Cited alongside, same era.
Faulty reward functions in the wild
Clark, J. and Amodei, D. (2016) · 2024
Closest in time.
Making a sota adversarial attack on llms 38x faster
Haize Labs (2024) · 2024
Closest in time.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J., Wu, X., Wang, W., Xianyu, Zhang, D., and Cao, Y. (2024) · 2024
Closest in time.
Preventing reward hacking with occupancy measure regularization
Laidlaw, C., Singhal, S., and Dragan, A. (2024) · 2024
Closest in time.
The alignment ceiling: Objective mismatch in reinforcement learning from human feedback
Lambert, N. and Calandra, R. (2024) · 2024
Closest in time.
Catastrophes, conspiracies, and subexponential distributions (part ii)
Wierman, A. (2013) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Starling-7B: Improving llm helpfulness & harmlessness with rlaif
Zhu, B., Frick, E., Wu, T., Zhu, H., and Jiao, J. (2023) · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. (2023) · 2023
Cited alongside, same era.
Closest in time.