Fetching the paper…
Reading the bibliography…
This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences.
Robust solutions of optimization problems affected by uncertain probabilities
Ben-Tal, A., den Hertog, D., De Waegenaere, A., Melenberg, B., and Rennen, G · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Willams, R. J · 1992
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K · 1999
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Offline learning in markov games with general function approximation
Zhang, Y., Bai, Y., and Jiang, N · 2003
Earlier work this paper cites.
The price of robustness
Bertsimas, D. and Sim, M · 2004
Earlier work this paper cites.
Fitted q-iteration in continuous action-space mdps
Antos, A., Szepesvári, C., and Munos, R · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A.-m., Szepesvári, C., and Munos, R · 2010
Earlier work this paper cites.
Individual Choice Behavior: A Theoretical Analysis
Luce, R · 2012
Earlier work this paper cites.
Concrete Problems in AI Safety, 2016
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
Völske, M., Potthast, M., Syed, S., and Stein, B · 2017
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Des Combes, R. T · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Earlier work this paper cites.
The importance of pessimism in fixed-dataset policy optimization
Buckman, J., Gelada, C., and Bellemare, M. G · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Provably good batch off-policy reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Later among the works it cites.
Nash Learning from Human Feedback, 2023
Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., Selvi, M., Girgin, S., Momchev, N., Bachem, O., Mankowitz, D. J., Precup, D., and Piot, B · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model, 2023
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Scaling up models and data with t5x and seqio
Roberts, A., Chung, H. W., Mishra, G., Levskaya, A., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., et al · 2023
Later among the works it cites.
A long way to go: Investigating length correlations in rlhf
Singhal, P., Goyal, T., Xu, J., and Durrett, G · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Cited alongside, same era.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., Foster, D. P., and Kakade, S. M · 2020
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Xie, T. and Jiang, N · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A · 2021
Cited alongside, same era.
Xiao, C., Wang, H., Pan, Y., White, A., and White, M · 2023
Later among the works it cites.
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs, 2024
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
Later among the works it cites.
Human Alignment of Large Language Models through Online Preference Optimisation, 2024
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., Joshi, R., Zheng, Z., and Piot, B · 2024
Later among the works it cites.
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF, 2024
Cen, S., Mei, J., Goshvadi, K., Dai, H., Yang, T., Yang, S., Schuurmans, D., Chi, Y., and Dai, B · 2024
Later among the works it cites.
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking, 2024
Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D. J., Fisch, A., Heller, K., Pfohl, S., Ramachandran, D., Shaw, P., and Berant, J · 2024
Later among the works it cites.
Robust preference optimization through reward model distillation
Fisch, A., Eisenstein, J., Zayats, V., Agarwal, A., Beirami, A., Nagpal, C., Shaw, P., and Berant, J · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team · 2024
Later among the works it cites.
Direct Language Model Alignment from Online AI Feedback, 2024
Guo, S., Zhang, B., Liu, T., Liu, T., Khalman, M., Llinares, F., Rame, A., Mesnard, T., Zhao, Y., Piot, B., Ferret, J., and Blondel, M · 2024
Later among the works it cites.
Information-directed pessimism for offline reinforcement learning
Koppel, A., Bhatt, S., Guo, J., Eappen, J., Wang, M., and Ganesh, S · 2024
Later among the works it cites.
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer, 2024
Liu, Z., Lu, M., Zhang, S., Liu, B., Guo, H., Yang, Y., Blanchet, J., and Wang, Z · 2024
Later among the works it cites.
Disentangling Length from Quality in Direct Preference Optimization, 2024
Park, R., Rafailov, R., Ermon, S., and Finn, C · 2024
Later among the works it cites.
A Minimaximalist Approach to Reinforcement Learning from Human Feedback, 2024
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A · 2024
Later among the works it cites.
Generalized Preference Optimization: A Unified Approach to Offline Alignment, 2024
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Pires, B. Á., and Piot, B · 2024
Later among the works it cites.
Online Iterative Reinforcement Learning from Human Feedback with General Preference Model, 2024
Ye, C., Xiong, W., Zhang, Y., Jiang, N., and Zhang, T · 2024
Later among the works it cites.
Pessimism meets risk: risk-sensitive offline reinforcement learning
Zhang, D., Lyu, B., Qiu, S., Kolar, M., and Zhang, T · 2024
Later among the works it cites.