Fetching the paper…
Reading the bibliography…
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Stimulus and response generalization: A stochastic model relating generalization to distance in psychological space
Roger N Shepard · 1957
Earlier work this paper cites.
Learning non-myopically from human-generated reward
W. Bradley Knox and Peter Stone · 1965
Earlier work this paper cites.
The Rating of Chessplayers, Past and Present
Arpad Elo · 1978
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 2005
Earlier work this paper cites.
Introducing machine learning within an interactive evolutionary design environment
AT Machwe and IC Parmee · 2006
Earlier work this paper cites.
Picbreeder: Evolving pictures collaboratively online
Jimmy Secretan, Nicholas Beato, David B D Ambrosio, Adelein Rodriguez, Adam Campbell, and Kenneth O Stanley · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
A bayesian interactive optimization approach to procedural animation design
Eric Brochu, Tyson Brochu, and Nando de Freitas · 2010
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
Patrick M Pilarski, Michael R Dawson, Thomas Degris, Farbod Fahimi, Jason P Carey, and Richard Sutton · 2011
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Riad Akrour, Marc Schoenauer, and Michèle Sebag · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: A formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Earlier work this paper cites.
Learning from human-generated reward
William Bradley Knox · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
A Bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Score-based inverse reinforcement learning
Layla El Asri, Bilal Piot, Matthieu Geist, Romain Laroche, and Olivier Pietquin · 2016
Later among the works it cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart Russell, Pieter Abbeel, and Anca Dragan · 2016
Later among the works it cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Later among the works it cites.
Active reinforcement learning: Observing rewards at a cost
David Krueger, Jan Leike, Owain Evans, and John Salvatier · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christian Wirth and Johannes Fürnkranz · 2013
Cited alongside, same era.
Programming by feedback
Riad Akrour, Marc Schoenauer, Michèle Sebag, and Jean-Christophe Souplet · 2014
Cited alongside, same era.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Cited alongside, same era.
Active reward learning
Christian Daniel, Malte Viering, Jan Metz, Oliver Kroemer, and Jan Peters · 2014
Cited alongside, same era.
Active reward learning with a novel acquisition function
Christian Daniel, Oliver Kroemer, Malte Viering, Jan Metz, and Jan Peters · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Later among the works it cites.
Should we fear supersmart robots?
Stuart Russell · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Breeding a diversity of super mario behaviors through interactive evolution
Patrikk D Sørensen, Jeppeh M Olsen, and Sebastian Risi · 2016
Later among the works it cites.
Learning language games through interaction
Sida I Wang, Percy Liang, and Christopher D Manning · 2016
Later among the works it cites.
Model-free preference-based reinforcement learning
Christian Wirth, J Fürnkranz, Gerhard Neumann, et al · 2016
Later among the works it cites.
Generalizing skills with semi-supervised reinforcement learning
Chelsea Finn, Tianhe Yu, Justin Fu, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Learning from demonstrations for real world reinforcement learning
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, Joel Z Leibo, and Audrunas Gruslys · 2017
Closest in time.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, David Roberts, Matthew E Taylor, and Michael L Littman · 2017
Closest in time.
Third-person imitation learning
Bradly C Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Closest in time.