Fetching the paper…
Reading the bibliography…
For imitation learning algorithms to scale to real-world challenges, they must handle high-dimensional observations, offline learning, and policy-induced covariate-shift.
Alvinn: An autonomous land vehicle in a neural network
Dean A. Pomerleau · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 2004
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 2005
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and J. Andrew Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charlie Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Learning from demonstrations for real world reinforcement learning
Todd Hester, Matej Vecerík, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John P. Agapiou, Joel Z. Leibo, and Audrunas Gruslys · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David R Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2019
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard S. Zemel, and Shixiang Shane Gu · 2019
Cited alongside, same era.
Rl baselines3 zoo
Antonin Raffin · 2020
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc G. Bellemare · 2021
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
Mobile: Model-based imitation learning from observation alone
Rahul Kidambi, Jonathan Chang, and Wen Sun · 2021
Later among the works it cites.
Visual adversarial imitation learning using variational models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning agile and dynamic motor skills for legged robots
Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, and Marco Hutter · 2019
Cited alongside, same era.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Cited alongside, same era.
Random expert distillation: Imitation learning via expert policy support estimation
Ruohan Wang, Carlo Ciliberto, Pierluigi Vito Amadori, and Y. Demiris · 2019
Cited alongside, same era.
Disagreement-regularized imitation learning
Kianté Brantley, Wen Sun, and Mikael Henaff · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy P. Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Cited alongside, same era.
Benchmarking end-to-end behavioural cloning on video games
Anssi Kanervisto, Joonas Pussinen, and Ville Hautamäki · 2020
Cited alongside, same era.
Rafael Rafailov, Tianhe Yu, Aravind Rajeswaran, and Chelsea Finn · 2021
Later among the works it cites.
Behavioral cloning from noisy demonstrations
Fumihiro Sasaki and Ryota Yamashina · 2021
Later among the works it cites.
Imitation learning by reinforcement learning
Kamil Ciosek · 2022
Later among the works it cites.
Demodice: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg, and P. Abbeel · 2022
Later among the works it cites.
Atari-head: Atari human eye-tracking and demonstration dataset
Ruohan Zhang, Zhuode Liu, L. Guan, Luxin Zhang, Mary M. Hayhoe, and Dana H. Ballard · 2022
Later among the works it cites.