Fetching the paper…
Reading the bibliography…
Preference-based reinforcement learning (PbRL) aligns a robot behavior with human preferences via a reward function learned from binary feedback over agent behaviors.
Rank analysis of incomplete block designs: I. The method of paired comparisons
R. Bradley and M. Terry · 1952
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W. Knox and P. Stone · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W. Knox and P. Stone · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Preference-based policy learning
R. Akrour, M. Schoenauer, and M. Sebag · 2011
Earlier work this paper cites.
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
P. Pilarski, M. Dawson, T. Degris, F. Fahimi, J. Carey, and R. Sutton · 2011
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
R. Akrour, M. Schoenauer, and M. Sebag · 2012
Earlier work this paper cites.
A Bayesian approach for policy learning from trajectory preference queries
A. Wilson, A. Fern, and P. Tadepalli · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
H. Sugiyama, T. Meguro, and Y. Minami · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: A preliminary survey
C. Wirth and J. Fürnkranz · 2013
Earlier work this paper cites.
ADAM: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. Russell, P. Abbeel, and A. Dragan · 2016
Earlier work this paper cites.
Model-free preference-based reinforcement learning
C. Wirth, J. Fürnkranz, and G. Neumann · 2016
Earlier work this paper cites.
Inverse reward design
D. Hadfield-Menell, S. Milli, P. Abbeel, S. Russell, and A. Dragan · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. Dragan, S. Sastry, and S. Seshia · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
J. MacGlashan, M. Ho, R. Loftin, B. Peng, G. Wang, D. Roberts, M. Taylor, and M. Littman · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in Atari
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: A research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Cited alongside, same era.
Deep Tamer: Interactive agent shaping in high-dimensional state spaces
G. Warnell, N. Waytowich, V. Lawhern, and P. Stone · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 2018
Cited alongside, same era.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning
E. Biyik, N. Huynh, M. Kochenderfer, and D. Sadigh · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Later among the works it cites.
Deep reinforcement and InfoMax learning
B. Mazoure, R. Tachet des Combes, T. Doan, P. Bachman, and R. Hjelm · 2020
Later among the works it cites.
Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Later among the works it cites.
PEBBLE: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
K. Lee, L. Smith, and P. Abbeel · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Deep reinforcement learning from policy-dependent human feedback
D. Arumugam, J. Lee, S. Saskin, and M. Littman · 2019
Cited alongside, same era.
Solar: Deep structured representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. Johnson, and S. Levine · 2019
Cited alongside, same era.
End-to-end robotic reinforcement learning without reward engineering
A. Singh, L. Yang, K. Hartikainen, C. Finn, and S. Levine · 2019
Cited alongside, same era.
AVID: Learning multi-stage tasks via pixel-level translation of human videos
L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S. Levine · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Cited alongside, same era.
J. Wu, L. Ouyang, D. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Later among the works it cites.
B-Pref: Benchmarking preference-based reinforcement learning
K. Lee, L. Smith, A. Dragan, and P. Abbeel · 2021
Later among the works it cites.
Skill preferences: Learning to extract and execute robotic skills from human feedback
X. Wang, K. Lee, K. Hakhamaneshi, P. Abbeel, and M. Laskin · 2021
Later among the works it cites.
Mastering Atari games with limited data
W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao · 2021
Later among the works it cites.
Learning from suboptimal demonstration via self-supervised reward regression
L. Chen, R. Paleja, and M. Gombolay · 2021
Later among the works it cites.
Exploring simple Siamese representation learning
X. Chen and K. He · 2021
Later among the works it cites.
B-Pref, 2021
K. Lee, L. Smith, A. Dragan, and P. Abbeel · 2021
Later among the works it cites.
Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards
P. Goyal, S. Niekum, and R. Mooney · 2021
Later among the works it cites.
J. Park, Y. Seo, J. Shin, H. Lee, P. Abbeel, and K. Lee · 2022
Later among the works it cites.
Meta-Reward-Net: Implicitly differentiable reward learning for preference-based reinforcement learning
R. Liu, F. Bai, Y. Du, and Y. Yang · 2022
Later among the works it cites.
Reward uncertainty for exploration in preference-based reinforcement learning
X. Liang, K. Shu, K. Lee, and P. Abbeel · 2022
Later among the works it cites.
A ranking game for imitation learning
H. Sikchi, A. Saran, W. Goo, and S. Niekum · 2022
Later among the works it cites.
Towards more generalizable one-shot visual imitation learning
Z. Mandi, F. Liu, K. Lee, and P. Abbeel · 2022
Later among the works it cites.
Sirl: Similarity-based implicit representation learning
A. Bobu, Y. Liu, R. Shah, D. S. Brown, and A. D. Dragan · 2023
Later among the works it cites.