Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
R. Akrour, M. Schoenauer, and M. Sebag · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
J. Fürnkranz, E. Hüllermeier, W. Cheng, and S.-H. Park · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
A. Wilson, A. Fern, and P. Tadepalli · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 2014
Earlier work this paper cites.
Variational inference: A review for statisticians
D. Blei, A. Kucukelbir, and J. McAuliffe · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Learning from physical human corrections, one feature at a time
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2018
Earlier work this paper cites.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Earlier work this paper cites.
Neurips 2018 demographics and inclusion survey: Summary of responses
H. Daume and K. Heller · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Earlier work this paper cites.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters · 2018
Earlier work this paper cites.
Asking easy questions: A user-friendly approach to active reward learning
E. Bıyık, M. Palan, N. C. Landolfi, D. P. Losey, and D. Sadigh · 2019
Earlier work this paper cites.
Multi-task deep reinforcement learning with popart
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. Van Hasselt · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Is more autonomy always better? exploring preferences of users with mobility impairments in robot-assisted feeding
T. Bhattacharjee, E. Gordon, et al · 2020
Earlier work this paper cites.
LESS is more: Rethinking probabilistic models of human behavior
A. Bobu, D. R. R. Scobee, J. F. Fisac, S. S. Sastry, and A. D. Dragan · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
I. Gabriel · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Uncertain decisions facilitate better preference learning
C. Laidlaw and S. Russell · 2021
Cited alongside, same era.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
The history and risks of reinforcement learning and human feedback
N. Lambert, T. K. Gilbert, and T. O. Zick · 2023
Later among the works it cites.
Nash learning from human feedback
R. Munos, M. Valko, D. Calandriello, M. G. Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, A. Michi, et al · 2023
Later among the works it cites.
Active reward learning from online preferences, 2023
V. Myers, E. Bıyık, and D. Sadigh · 2023
Later among the works it cites.
Distributional preference learning: Understanding and accounting for hidden context in rlhf
A. Siththaranjan, C. Laidlaw, and D. Hadfield-Menell · 2023
Later among the works it cites.
Breadcrumbs to the goal: Goal-conditioned exploration from human-in-the-loop feedback, 2023
M. Torne, M. Balsells, Z. Wang, S. Desai, T. Chen, P. Agrawal, and A. Gupta · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
E. Bıyık, D. P. Losey, M. Palan, N. C. Landolfi, G. Shevchuk, and D. Sadigh · 2022
Cited alongside, same era.
vec2text with round-trip translations
G. Cideron, S. Girgin, A. Raichuk, O. Pietquin, O. Bachem, and L. Hussenot · 2022
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Cited alongside, same era.
The boltzmann policy distribution: Accounting for systematic suboptimality in human models
C. Laidlaw and A. Dragan · 2022
Cited alongside, same era.
Goal-conditioned reinforcement learning: Problems and solutions
M. Liu, M. Zhu, and W. Zhang · 2022
Cited alongside, same era.
Sgpt: Gpt sentence embeddings for semantic search
N. Muennighoff · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Is rlhf more difficult than standard rl? a theoretical perspective
Y. Wang, Q. Liu, and C. Jin · 2023
Later among the works it cites.
Habitat 3.0: A co-habitat for humans, avatars and robots, 2023
P. Xavi, U. Eric, A. Szot, M. D. Cote, R. Partsey, J. Yang, R. Desai, A. W. Clegg, M. Hlavac, T. Min, T. Gervet, V. Vondrus, V.-P. Berges, J. Turner, O. Maksymets, Z. Kira, M. Kalakrishnan, J. Malik, D. S. Chaplot, U. Jain, D. Batra, A. Rai, and R. Mottaghi · 2023
Later among the works it cites.
Group preference optimization: Few-shot alignment of large language models
S. Zhao, J. Dang, and A. Grover · 2023
Later among the works it cites.
Beyond one-preference-for-all: Multi-objective direct preference optimization
Z. Zhou, J. Liu, C. Yang, J. Shao, Y. Liu, X. Yue, W. Ouyang, and Y. Qiao · 2023
Later among the works it cites.
LLM2Vec: Large language models are secretly powerful text encoders
P. BehnamGhader, V. Adlakha, M. Mosbach, D. Bahdanau, N. Chapados, and S. Reddy · 2024
Closest in time.
Pareto-optimal learning from preferences with hidden context
R. Boldi, L. Ding, L. Spector, and S. Niekum · 2024
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
S. Chakraborty, J. Qiu, H. Yuan, A. Koppel, F. Huang, D. Manocha, A. S. Bedi, and M. Wang · 2024
Closest in time.
Social choice for ai alignment: Dealing with diverse human feedback
V. Conitzer, R. Freedman, J. Heitzig, W. H. Holliday, B. M. Jacobs, N. Lambert, M. Mossé, E. Pacuit, S. Russell, H. Schoelkopf, et al · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
Modular pluralism: Pluralistic alignment via multi-llm collaboration
S. Feng, T. Sorensen, Y. Liu, J. R. Fisher, C. Y. Park, Y. Choi, and Y. Tsvetkov · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang · 2024
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate
K. Kang, E. Wallace, C. Tomlin, A. Kumar, and S. Levine · 2024
Closest in time.
H. R. Kirk, A. Whitefield, P. Röttger, A. Bean, K. Margatina, J. Ciro, R. Mosquera, M. Bartolo, A. Williams, H. He, et al · 2024
Closest in time.
Personalized language modeling from personalized human feedback
X. Li, Z. C. Lipton, and L. Leqi · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
A roadmap to pluralistic alignment
T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, et al · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
G. Swamy, C. Dann, R. Kidambi, Z. S. Wu, and A. Agarwal · 2024
Closest in time.
Self-exploring language models: Active preference elicitation for online alignment
S. Zhang, D. Yu, H. Sharma, Z. Yang, S. Wang, H. Hassan, and Z. Wang · 2024
Closest in time.