Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) has been crucial to the recent success of Large Language Models (LLMs), however, it is often a complex and brittle process.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
The ‘awful idea of accountability’: inscribing people into the measurement of objects
K. Hoskin · 1996
Earlier work this paper cites.
Maximum margin planning
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W. B. Knox and P. Stone · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Preference-based policy learning
R. Akrour, M. Schoenauer, and M. Sebag · 2011
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Faulty reward functions in the wild, 2016
J. Clark and D. Amodei · 2016
Earlier work this paper cites.
Quantilizers: A safer alternative to maximizers for limited optimization
J. Taylor · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
T. Everitt, V. Krakovna, L. Orseau, M. Hutter, and S. Legg · 2017
Earlier work this paper cites.
Inverse reward design
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan · 2017
Earlier work this paper cites.
Tactics of adversarial attack on deep reinforcement learning agents
Y.-C. Lin, Z.-W. Hong, Y.-H. Liao, M.-L. Shih, M.-Y. Liu, and M. Sun · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
On adversarial examples for character-level neural machine translation
J. Ebrahimi, D. Lowd, and D. Dou · 2018
Earlier work this paper cites.
Generalization and regularization in dqn
J. Farebrother, M. C. Machado, and M. Bowling · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration, 2019
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Classifying specification problems as variants of goodhart’s law, 8 2019
V. Krakovna and R. Kumar · 2019
Earlier work this paper cites.
Categorizing variants of goodhart’s law, 2019
D. Manheim and S. Garrabrant · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models, 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Fine-tuning language models from human preferences, 2020
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2020
Cited alongside, same era.
A long way to go: Investigating length correlations in rlhf, 2023
P. Singhal, T. Goyal, J. Xu, and G. Durrett · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization, 2023
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2023
Later among the works it cites.
Coherent soft imitation learning
J. Watson, S. Huang, and N. Heess · 2023
Later among the works it cites.
Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, 2023
Y. Zhai, H. Zhang, Y. Lei, Y. Yu, K. Xu, D. Feng, B. Ding, and H. Wang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
A. Pan, K. Bhatia, and J. Steinhardt · 2022
Cited alongside, same era.
A ranking game for imitation learning
H. Sikchi, A. Saran, W. Goo, and S. Niekum · 2022
Cited alongside, same era.
Defining and characterizing reward hacking, 2022
J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger · 2022
Cited alongside, same era.
Learning to summarize from human feedback, 2022
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano · 2022
Cited alongside, same era.
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, A. Üstün, and S. Hooker · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback, 2024
Y. Dubois, X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. Liang, and T. B. Hashimoto · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
Contrastive preference learning: Learning from human feedback without reinforcement learning
J. Hejna, R. Rafailov, H. Sikchi, C. Finn, S. Niekum, W. B. Knox, and D. Sadigh · 2024
Closest in time.
V-star: Training verifiers for self-taught reasoners
A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal · 2024
Closest in time.
Understanding the learning dynamics of alignment with human feedback, 2024
S. Im and Y. Li · 2024
Closest in time.
Mixtral of experts, 2024
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
Statistical rejection sampling improves preference optimization, 2024
T. Liu, Y. Zhao, R. Joshi, M. Khalman, M. Saleh, P. J. Liu, and J. Liu · 2024
Closest in time.
Confronting reward model overoptimization with constrained RLHF
T. Moskovitz, A. K. Singh, D. Strouse, T. Sandholm, R. Salakhutdinov, A. Dragan, and S. M. McAleer · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
A. Pal, D. Karkhanis, S. Dooley, M. Roberts, S. Naidu, and C. White · 2024
Closest in time.
Disentangling length from quality in direct preference optimization, 2024
R. Park, R. Rafailov, S. Ermon, and C. Finn · 2024
Closest in time.
From r r to q ∗ q^{*} : Your language model is secretly a q-function, 2024
R. Rafailov, J. Hejna, R. Park, and C. Finn · 2024
Closest in time.
Countering reward over-optimization in llm with demonstration-guided reinforcement learning
M. Rita, F. Strub, R. Chaabouni, P. Michel, E. Dupoux, and O. Pietquin · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
F. Tajwar, A. Singh, A. Sharma, R. Rafailov, J. Schneider, T. Xie, S. Ermon, C. Finn, and A. Kumar · 2024
Closest in time.
Iterative reasoning preference optimization
R. Yuanzhe Pang, W. Yuan, K. Cho, H. He, S. Sukhbaatar, and J. Weston · 2024
Closest in time.
Iterative data smoothing: Mitigating reward overfitting and overoptimization in rlhf
B. Zhu, M. I. Jordan, and J. Jiao · 2024
Closest in time.