Fetching the paper…
Reading the bibliography…
In this paper, we propose to combine imitation and reinforcement learning via the idea of reward shaping using an oracle.
Learning to predict by the methods of temporal differences
RichardS. Sutton · 1988
Earlier work this paper cites.
Subgrouping reduces complexity and speeds up learning in recurrent networks
David Zipser · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Shaping and policy search in reinforcement learning
Andrew Y Ng · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Policy search by dynamic programming
J Andrew Bagnell, Sham M Kakade, Jeff G Schneider, and Andrew Y Ng · 2004
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé III, John Langford, and Daniel Marcu · 2009
Cited alongside, same era.
Classification-based policy iteration with a critic
Victor Gabillon, Alessandro Lazaric, Mohammad Ghavamzadeh, and Bruno Scherrer · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J Gordon, and J.Andrew Bagnell · 2011
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Cited alongside, same era.
Learning to search better than your teacher
Kai-wei Chang, Akshay Krishnamurthy, Alekh Agarwal, Hal Daume, and John Langford · 2015
Cited alongside, same era.
Truncated approximate dynamic programming with task-dependent terminal value
Amir-massoud Farahmand, Daniel Nikolaev Nikovski, Yuji Igarashi, and Hiroki Konaka · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver et al · 2016
Later among the works it cites.
Learning to filter with predictive state inference machines
Wen Sun, Arun Venkatraman, Byron Boots, and J Andrew Bagnell · 2016
Later among the works it cites.
Learning to gather information via imitation
Sanjiban Choudhury, Ashish Kapoor, Gireeja Ranade, and Debadeepta Dey · 2017
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Agile off-road autonomous driving using end-to-end deep imitation learning
Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos Theodorou, and Byron Boots · 2017
Later among the works it cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, John Schulman, Emanuel Todorov, and Sergey Levine · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J Gordon, Byron Boots, and J Andrew Bagnell · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Večerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Later among the works it cites.