Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) commonly assumes access to well-specified reward functions, which many practical applications do not provide.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
A user-centric evaluation framework for recommender systems
Pearl Pu, Li Chen, and Rong Hu · 2011
Earlier work this paper cites.
Personality-based recommender systems: an overview
Maria Augusta SN Nunes and Rong Hu · 2012
Earlier work this paper cites.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Human decision making and recommender systems
Anthony Jameson, Martijn C Willemsen, Alexander Felfernig, Marco de Gemmis, Pasquale Lops, Giovanni Semeraro, and Li Chen · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Learning robot objectives from physical human interaction
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
The off-switch game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Cited alongside, same era.
Imitation learning: A survey of learning methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al · 2017
Cited alongside, same era.
A review of user interface design for interactive machine learning
John J Dudley and Per Ola Kristensson · 2018
Cited alongside, same era.
Learning to walk via deep reinforcement learning
Tuomas Haarnoja, Sehoon Ha, Aurick Zhou, Jie Tan, George Tucker, and Sergey Levine · 2018
Cited alongside, same era.
Human-centered explainable ai: towards a reflective sociotechnical approach
Upol Ehsan and Mark O Riedl · 2020
Later among the works it cites.
Choice set misspecification in reward inference
Rachel Freedman, Rohin Shah, and Anca Dragan · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Later among the works it cites.
A review on interactive reinforcement learning from human social feedback
Jinying Lin, Zhen Ma, Randy Gomez, Keisuke Nakamura, Bo He, and Guangliang Li · 2020
Later among the works it cites.
The who in explainable ai: how ai background shapes perceptions of ai explanations
Upol Ehsan, Samir Passi, Q Vera Liao, Larry Chan, I Lee, Michael Muller, Mark O Riedl, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qian Zhao, F Maxwell Harper, Gediminas Adomavicius, and Joseph A Konstan · 2018
Cited alongside, same era.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
On the feasibility of learning, rather than assuming, human biases for reward inference
Rohin Shah, Noah Gundotra, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
Preferences implicit in the state of the world
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
The empathic framework for task learning from implicit human feedback
Yuchen Cui, Qiping Zhang, Alessandro Allievi, Peter Stone, Scott Niekum, and W Bradley Knox · 2020
Cited alongside, same era.
W Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2021
Later among the works it cites.
B-Pref: Benchmarking preference-based reinforcement learning
Kimin Lee, Laura Smith, Anca Dragan, and Pieter Abbeel · 2021
Later among the works it cites.
A survey of human-centered evaluations in human-centered machine learning
Fabian Sperrle, Mennatallah El-Assady, Grace Guo, Rita Borgo, D Horng Chau, Alex Endert, and Daniel Keim · 2021
Later among the works it cites.
Co-adaptive visual data analysis and guidance processes
Fabian Sperrle, Astrik Jeitler, Jürgen Bernard, Daniel Keim, and Mennatallah El-Assady · 2021
Later among the works it cites.