Fetching the paper…
Reading the bibliography…
Human explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning.
Atari-head: Atari human eye-tracking and demonstration dataset
Ruohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan, Karl S Muller, Jake A Whritner, Luxin Zhang, Mary M Hayhoe, and Dana H Ballard · 1903
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Andrea Lockerd Thomaz, Cynthia Breazeal, et al · 2006
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Combining manual feedback with subsequent mdp reward signals for reinforcement learning
W Bradley Knox and Peter Stone · 2010
Earlier work this paper cites.
Reinforcement learning from simultaneous human and mdp reward
W Bradley Knox and Peter Stone · 2012
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz · 2013
Earlier work this paper cites.
Learning something from nothing: Leveraging implicit human feedback strategies
Robert Loftin, Bei Peng, James MacGlashan, Michael L Littman, Matthew E Taylor, Jeff Huang, and David L Roberts · 2014
Earlier work this paper cites.
Boosted bellman residual minimization handling expert demonstrations
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Earlier work this paper cites.
Policy shaping with human teachers
Thomas Cederborg, Ishaan Grover, Charles L Isbell, and Andrea L Thomaz · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Learning from explanations using sentiment and advice in rl
Samantha Krening, Brent Harrison, Karen M Feigh, Charles Lee Isbell, Mark Riedl, and Andrea Thomaz · 2016
Earlier work this paper cites.
Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
Robert Loftin, Bei Peng, James MacGlashan, Michael L Littman, Matthew E Taylor, Jeff Huang, and David L Roberts · 2016
Earlier work this paper cites.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Visualizing and understanding atari agents
Sam Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman · 2017
Cited alongside, same era.
Right for the right reasons: Training differentiable models by constraining their explanations
Explanatory interactive machine learning
Stefano Teso and Kristian Kersting · 2019
Later among the works it cites.
Leveraging human guidance for deep reinforcement learning tasks
Ruohan Zhang, Faraz Torabi, Lin Guan, Dana H. Ballard, and Peter Stone · 2019
Later among the works it cites.
The empathic framework for task learning from implicit human feedback
Yuchen Cui, Qiping Zhang, Alessandro Allievi, Peter Stone, Scott Niekum, and W Bradley Knox · 2020
Closest in time.
Using human gaze to improve robustness against irrelevant objects in robot manipulation tasks
Heecheol Kim, Yoshiyuki Ohmura, and Yasuo Kuniyoshi · 2020
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Closest in time.
Reinforcement learning with augmented data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Cited alongside, same era.
Dqn-tamer: Human-in-the-loop reinforcement learning with intractable feedback
Riku Arakawa, Sosuke Kobayashi, Yuya Unno, Yuta Tsuboi, and Shin-ichi Maeda · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2018
Cited alongside, same era.
Bridging the gap: Converting human advice into imagined examples
Eric Yeh, Melinda Gervasio, Daniel Sanchez, Matthew Crossley, and Karen Myers · 2018
Cited alongside, same era.
Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Closest in time.
Representation learning via invariant causal mechanisms
Jovana Mitrovic, Brian McWilliams, Jacob Walker, Lars Buesing, and Charles Blundell · 2020
Closest in time.
Demystifying contrastive self-supervised learning: Invariances, augmentations and dataset biases
Senthil Purushwalkam and Abhinav Gupta · 2020
Closest in time.
Automatic data augmentation for generalization in deep reinforcement learning
Roberta Raileanu, Max Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2020
Closest in time.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
Laura Rieger, Chandan Singh, William Murdoch, and Bin Yu · 2020
Closest in time.
Efficiently guiding imitation learning algorithms with human gaze
Akanksha Saran, Ruohan Zhang, Elaine Schaertl Short, and Scott Niekum · 2020
Closest in time.
Making deep neural networks right for the right scientific reasons by interacting with their explanations
Patrick Schramowski, Wolfgang Stammer, Stefano Teso, Anna Brugger, Franziska Herbert, Xiaoting Shao, Hans-Georg Luigs, Anne-Katrin Mahlein, and Kristian Kersting · 2020
Closest in time.
Fresh: Interactive reward shaping in high-dimensional state spaces using human feedback
Baicen Xiao, Qifan Lu, Bhaskar Ramasubramanian, Andrew Clark, Linda Bushnell, and Radha Poovendran · 2020
Closest in time.
Human gaze assisted artificial intelligence: A review
R Zhang, A Saran, B Liu, Y Zhu, S Guo, S Niekum, D Ballard, and M Hayhoe · 2020
Closest in time.
Feature expansive reward learning: Rethinking human input
Andreea Bobu, Marius Wiggert, Claire Tomlin, and Anca D Dragan · 2021
Closest in time.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Kimin Lee, Laura Smith, and Pieter Abbeel · 2021
Closest in time.
Symbols as a lingua franca for bridging human-ai chasm for explainable and advisable ai systems
Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, and Lin Guan · 2022
Closest in time.