Fetching the paper…
Reading the bibliography…
Despite its promise, reinforcement learning's real-world adoption has been hampered by the need for costly exploration to learn a good policy.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, and Jan Peters · 1935
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
ALVINN: An autonomous land vehicle in a neural network
Dean Pomerleau · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas, Dimitri P Bertsekas, Dimitri P Bertsekas, and Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Peter Abbeel and Andrew Ng · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J. Andrew Bagnell, and Jan Peters · 2013
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Learning to search better than your teacher
Reinforcement learning on web interfaces using workflow-guided exploration
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang · 2018
Later among the works it cites.
Simple random search of static linear policies is competitive for reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Later among the works it cites.
Fast policy learning through imitation and reinforcemen
Ching-An Cheng, Xinyan Yan, Nolan Wagener, and Byron Boots · 2018
Later among the works it cites.
OIL: Observational imitation learning
Guohao Li, Matthias Mueller, Vincent Casser, Neil Smith, Dominik L Michels, and Bernard Ghanem · 2018
Later among the works it cites.
Convergence of value aggregation for imitation learning
Ching-An Cheng and Byron Boots · 2018
Later among the works it cites.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kai-Wei Chang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Hal Daumé III · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn and Sergey Levine · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Cited alongside, same era.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J Gordon, Byron Boots, and J Andrew Bagnell · 2017
Cited alongside, same era.
Infogail: Interpretable imitation learning from visual demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon · 2017
Cited alongside, same era.
Wen Sun, J Andrew Bagnell, and Byron Boots · 2018
Later among the works it cites.
Dart: Dynamic animation and robotics toolkit
Jeongseok Lee, Michael Grey, Sehoon Ha, Tobias Kunz, Sumit Jain, Yuting Ye, Siddhartha Srinivasa, Mike Stilman, and Chuanjian Liu · 2018
Later among the works it cites.
Imitation learning from visual data with multiple intentions
Aviv Tamar, Khashayar Rohanimanesh, Yinlam Chow, Chris Vigorito, Ben Goodrich, Michael Kahane, and Derik Pridmore · 2018
Later among the works it cites.
Reinforcement learning with multiple experts: A Bayesian model combination approach
Michael Gimelfarb, Scott Sanner, and Chi-Guhn Lee · 2018
Later among the works it cites.
Applications of deep reinforcement learning in communications and networking: A survey
Nguyen Cong Luong, Dinh Thai Hoang, Shimin Gong, Dusit Niyato, Ping Wang, Ying-Chang Liang, and Dong In Kim · 2019
Later among the works it cites.
Managing fog networks using reinforcement learning based load balancing algorithm
Jung-Yeon Baek, Georges Kaddoum, Sahil Garg, Kuljeet Kaur, and Vivianne Gravel · 2019
Later among the works it cites.
AC-Teach: A Bayesian actor-critic method for policy learning with an ensemble of suboptimal teachers
Andrey Kurenkov, Ajay Mandlekar, Roberto Martin-Martin, Silvio Savarese, and Animesh Garg · 2019
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
André Barreto, Shaobo Hou, Diana Borsa, David Silver, and Doina Precup · 2020
Closest in time.