Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) has emerged as a popular paradigm for aligning models with human intent.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Approximate gradient methods in policy-space optimization of markov reward processes
Peter Marbach and John N Tsitsiklis · 2003
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Riad Akrour, Marc Schoenauer, and Michèle Sebag · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Active reward learning with a novel acquisition function
Christian Daniel, Oliver Kroemer, Malte Viering, Jan Metz, and Jan Peters · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Deep reinforcement learning and the deadly triad
Hado Van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Earlier work this paper cites.
Learning to extract coherent summary via deep reinforcement learning
Yuxiang Wu and Baotian Hu · 2018
Cited alongside, same era.
The green choice: Learning and influencing human decisions on shared roads
Erdem Bıyık, Daniel A Lazar, Dorsa Sadigh, and Ramtin Pedarsani · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning
Erdem Biyik, Nicolas Huynh, Mykel J. Kochenderfer, and Dorsa Sadigh · 2020
Cited alongside, same era.
Offline preference-based apprenticeship learning
Daniel Shin and Daniel S Brown · 2021
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
Benjamin Eysenbach, Tianjun Zhang, Sergey Levine, and Russ R Salakhutdinov · 2022
Later among the works it cites.
Few-shot preference learning for human-in-the-loop RL
Joey Hejna and Dorsa Sadigh · 2022
Later among the works it cites.
Models of human preference for learning reward functions
W Bradley Knox, Stephane Hatgis-Kessell, Serena Booth, Scott Niekum, Peter Stone, and Alessandro Allievi · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Safe imitation learning via fast bayesian reward inference from preferences
Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
Reinforcement learning with augmented data
Misha Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2020
Cited alongside, same era.
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2022
Later among the works it cites.
Direct preference-based policy optimization without reward modeling
Gaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka, Kyung-Min Kim, and Hyun Oh Song · 2023
Closest in time.
Training diffusion models with reinforcement learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine · 2023
Closest in time.
Extreme q-learning: Maxent RL without entropy
Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon · 2023
Closest in time.
Inverse preference learning: Preference-based rl without a reward function
Joey Hejna and Dorsa Sadigh · 2023
Closest in time.
Distance weighted supervised learning for offline interaction data
Joey Hejna, Jensen Gao, and Dorsa Sadigh · 2023
Closest in time.
Beyond reward: Offline preference-guided policy optimization
Yachen Kang, Diyuan Shi, Jinxin Liu, Li He, and Donglin Wang · 2023
Closest in time.
Preference transformer: Modeling human preferences using transformers for rl
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2023
Closest in time.
Learning optimal advantage from preferences and mistaking it for reward, 2023
W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson, Serena Booth, Anca Dragan, Peter Stone, and Scott Niekum · 2023
Closest in time.
Aligning text-to-image models using human feedback
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Closest in time.
VIP: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Shattering the agent-environment interface for fine-tuning inclusive language models
Wanqiao Xu, Shi Dong, Dilip Arumugam, and Benjamin Van Roy · 2023
Closest in time.