Fetching the paper…
Reading the bibliography…
For many tasks, the reward function is inaccessible to introspection or too complex to be specified procedurally, and must instead be learned from user data.
An upper bound on the loss from approximate optimal-value functions
Satinder P. Singh and Richard C. Yee · 1994
Earlier work this paper cites.
Solving Least Squares Problems
Charles L. Lawson and Richard J. Hanson · 1995
Earlier work this paper cites.
Policy invariance under reward transformations: theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart Russell · 2000
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
A Bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Learning from human preferences, June 2017
Dario Amodei, Paul Christiano, and Alex Ray · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D. Dragan, S. Shankar Sastry, and Sanjit A. Seshia · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Learning to understand goal specifications by modelling reward
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Arian Hosseini, Pushmeet Kohli, and Edward Grefenstette · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Later among the works it cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, Oleg Sushkov, David Barker, Jonathan Scholz, Misha Denil, Nando de Freitas, and Ziyu Wang · 2019
Later among the works it cites.
Solving Rubik’s Cube with a robot hand
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 2019
Later among the works it cites.
A practical approach to insertion with variable socket position using deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning robust rewards with adverserial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Cited alongside, same era.
Stable Baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
OpenAI Five
OpenAI · 2018
Cited alongside, same era.
Mel Vecerik, Oleg Sushkov, David Barker, Thomas Rothörl, Todd Hester, and Jon Scholz · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander S. Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L. Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver · 2019
Later among the works it cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Later among the works it cites.
Understanding learned reward functions
Eric J. Michaud, Adam Gleave, and Stuart Russell · 2020
Closest in time.
imitation: implementations of inverse reinforcement learning and imitation learning algorithms
Steven Wang, Adam Gleave, and Sam Toyer · 2020
Closest in time.