Fetching the paper…
Reading the bibliography…
Preference-based reinforcement learning (PbRL) provides a natural way to align RL agents' behavior with human desired outcomes, but is often restrained by costly human feedback.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Clustering by means of medoids
LKPJ Rdusseeun and P Kaufman · 1987
Earlier work this paper cites.
Locally weighted learning for control
Christopher G Atkeson, Andrew W Moore, and Stefan Schaal · 1997
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J Andrew Bagnell and Jeff G Schneider · 2001
Earlier work this paper cites.
Fast poisson disk sampling in arbitrary dimensions
Robert Bridson · 2007
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
Patrick M Pilarski, Michael R Dawson, Thomas Degris, Farbod Fahimi, Jason P Carey, and Richard S Sutton · 2011
Earlier work this paper cites.
The optimal reward problem: Designing effective reward for bounded agents
Jonathan Daniel Sorg · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee et al · 2013
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Earlier work this paper cites.
Sample-based informationl-theoretic stochastic optimal control
Rudolf Lioutikov, Alexandros Paraschos, Jan Peters, and Gerhard Neumann · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Learning contact-rich manipulation skills with guided policy search
Sergey Levine, Nolan Wagener, and Pieter Abbeel · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
One-shot learning of manipulation skills with online dynamics adaptation and neural network priors
Justin Fu, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Cited alongside, same era.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
A deeper look at experience replay
Shangtong Zhang and Richard S Sutton · 2017
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
On the expressivity of markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K Ho, Michael Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Regret minimization experience replay in off-policy reinforcement learning
Xu-Hui Liu, Zhenghai Xue, Jingcheng Pang, Shengyi Jiang, Feng Xu, and Yang Yu · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch active preference-based learning of reward functions
Erdem Biyik and Dorsa Sadigh · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Universal planning networks: Learning generalizable representations for visuomotor control
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning
Erdem Biyik, Nicolas Huynh, Mykel Kochenderfer, and Dorsa Sadigh · 2020
Cited alongside, same era.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Later among the works it cites.
Reward uncertainty for exploration in preference-based reinforcement learning
Xinran Liang, Katherine Shu, Kimin Lee, and Pieter Abbeel · 2022
Later among the works it cites.
Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Runze Liu, Fengshuo Bai, Yali Du, and Yaodong Yang · 2022
Later among the works it cites.
SURF: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning
Jongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2022
Later among the works it cites.
Experience replay with likelihood-free importance weights
Samarth Sinha, Jiaming Song, Animesh Garg, and Stefano Ermon · 2022
Later among the works it cites.
Mind the gap: Offline policy optimization for imperfect rewards
Jianxiong Li, Xiao Hu, Haoran Xu, Jingjing Liu, Xianyuan Zhan, Qing-Shan Jia, and Ya-Qin Zhang · 2023
Closest in time.
Learning policy-aware models for model-based reinforcement learning via transition occupancy matching
Yecheng Jason Ma, Kausik Sivakumar, Jason Yan, Osbert Bastani, and Dinesh Jayaraman · 2023
Closest in time.
Benchmarks and algorithms for offline preference-based reward learning
Daniel Shin, Anca Dragan, and Daniel S. Brown · 2023
Closest in time.
Causal confusion and reward misidentification in preference-based reward learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca Dragan, and Daniel S Brown · 2023
Closest in time.
Live in the moment: Learning dynamics model adapted to evolving policy
Xiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, and Furong Huang · 2023
Closest in time.