Fetching the paper…
Reading the bibliography…
Reinforcement learning from Human Feedback (RLHF) learns from preference signals, while standard Reinforcement Learning (RL) directly learns from reward signals.
Individual choice behavior: A theoretical analysis
Robert D Luce · 1959
Earlier work this paper cites.
Aggregation of preference orderings
Germain Kreweras · 1965
Earlier work this paper cites.
Intransitivity of preferences
Amos Tversky · 1969
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Probabilistic social choice based on simple voting comparisons
Peter C Fishburn · 1984
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
David R Hunter · 2004
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yisong Yue and Thorsten Joachims · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims · 2012
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Reducing dueling bandits to cardinal bandits
Nir Ailon, Zohar Karnin, and Thorsten Joachims · 2014
Earlier work this paper cites.
Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
Róbert Busa-Fekete, Balázs Szörényi, Paul Weng, Weiwei Cheng, and Eyke Hüllermeier · 2014
Earlier work this paper cites.
Relative upper confidence bound for the k-armed dueling bandit problem
Masrour Zoghi, Shimon Whiteson, Remi Munos, and Maarten Rijke · 2014
Earlier work this paper cites.
Contextual dueling bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Consistent probabilistic social choice
Florian Brandl, Felix Brandt, and Hans Georg Seedig · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Preference-based reinforcement learning with finite-time guarantees
Yichong Xu, Ruosong Wang, Lin Yang, Aarti Singh, and Artur Dubrawski · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Preference-based online learning with dueling bandits: A survey
Viktor Bengs, Róbert Busa-Fekete, Adil El Mesaoudi-Paul, and Eyke Hüllermeier · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dqn-tamer: Human-in-the-loop reinforcement learning with intractable feedback
Riku Arakawa, Sosuke Kobayashi, Yuya Unno, Yuta Tsuboi, and Shin-ichi Maeda · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Learning adversarial mdps with bandit feedback and unknown transition
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 2019
Cited alongside, same era.
Online learning to rank for sequential music recommendation
Bruno L Pereira, Alberto Ueda, Gustavo Penha, Rodrygo LT Santos, and Nivio Ziviani · 2019
Cited alongside, same era.
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2021
Later among the works it cites.
Improving multimodal interactive agents with reinforcement learning from human feedback
Josh Abramson, Arun Ahuja, Federico Carnevale, Petko Georgiev, Alex Goldin, Alden Hung, Jessica Landon, Jirka Lhotka, Timothy Lillicrap, Alistair Muldal, et al · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
A posterior sampling framework for interactive decision making
Han Zhong, Wei Xiong, Sirui Zheng, Liwei Wang, Zhaoran Wang, Zhuoran Yang, and Tong Zhang · 2022
Later among the works it cites.
Learning a universal human prior for dexterous manipulation from human preference
Zihan Ding, Yuanpei Chen, Allen Z Ren, Shixiang Shane Gu, Hao Dong, and Chi Jin · 2023
Closest in time.
Aligning text-to-image models using human feedback
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Closest in time.
Improved regret for efficient online reinforcement learning with linear function approximation
Uri Sherman, Tomer Koren, and Yishay Mansour · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Banghua Zhu, Jiantao Jiao, and Michael I Jordan · 2023
Closest in time.