Fetching the paper…
Reading the bibliography…
Reward functions are difficult to design and often hard to align with human intent.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Rewarding behaviors
Fahiem Bacchus, Craig Boutilier, and Adam Grove · 1996
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W Bradley Knox and Peter Stone · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
Keyframe-based learning from demonstration
Baris Akgun, Maya Cakmak, Karl Jiang, and Andrea L Thomaz · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Active reward learning with a novel acquisition function
Christian Daniel, Oliver Kroemer, Malte Viering, Jan Metz, and Jan Peters · 2015
Earlier work this paper cites.
Data-driven motion mappings improve transparency in teleoperation
Rebecca P Khurshid and Katherine J Kuchenbecker · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Do you want your autonomous car to drive like you?
Chandrayee Basu, Qian Yang, David Hungerman, Mukesh Sinahal, and Anca D Draqan · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
Visual closed-loop control for pouring liquids
C. Schenck and D. Fox · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
What matters in learning from offline human demonstrations for robot manipulation
Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Martín-Martín · 2021
Later among the works it cites.
{AWAC}: Accelerating online reinforcement learning with offline datasets, 2021
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Later among the works it cites.
Offline preference-based apprenticeship learning
Daniel Shin and Daniel S Brown · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
The green choice: Learning and influencing human decisions on shared roads
Erdem Bıyık, Daniel A Lazar, Dorsa Sadigh, and Ramtin Pedarsani · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning
Erdem Biyik, Nicolas Huynh, Mykel J. Kochenderfer, and Dorsa Sadigh · 2020
Cited alongside, same era.
Safe imitation learning via fast bayesian reward inference from preferences
Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum · 2020
Cited alongside, same era.
Implementation matters in deep rl: A case study on ppo and trpo
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2020
Cited alongside, same era.
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nissan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Later among the works it cites.
Non-markovian reward modelling from trajectory labels via interpretable multiple instance learning
Joseph Early, Tom Bewley, Christine Evers, and Sarvapali Ramchurn · 2022
Later among the works it cites.
Few-shot preference learning for human-in-the-loop RL
Joey Hejna and Dorsa Sadigh · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Later among the works it cites.
Inferring rewards from language in context
Jessy Lin, Daniel Fried, Dan Klein, and Anca Dragan · 2022
Later among the works it cites.
Learning multimodal rewards from rankings
Vivek Myers, Erdem Biyik, Nima Anari, and Dorsa Sadigh · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning
Jongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2022
Later among the works it cites.
LS-IQ: Implicit reward regularization for inverse reinforcement learning
Firas Al-Hafez, Davide Tateo, Oleg Arenz, Guoping Zhao, and Jan Peters · 2023
Closest in time.
Extreme q-learning: Maxent RL without entropy
Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon · 2023
Closest in time.
Contrastive preference learning: Learning from human feedback without rl
Joey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn, Scott Niekum, W Bradley Knox, and Dorsa Sadigh · 2023
Closest in time.
Beyond reward: Offline preference-guided policy optimization
Yachen Kang, Diyuan Shi, Jinxin Liu, Li He, and Donglin Wang · 2023
Closest in time.
Preference transformer: Modeling human preferences using transformers for rl
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Offline rl with no ood actions: In-sample learning via implicit value regularization
Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Victor Wai Kin Chan, and Xianyuan Zhan · 2023
Closest in time.