Fetching the paper…
Reading the bibliography…
Reward function specification, which requires considerable human effort and iteration, remains a major impediment for learning behaviors through deep reinforcement learning.
. a general class of coefficients of divergence of one distribution from another
Syed Mumtaz Ali and Samuel Silvey · 1966
Earlier work this paper cites.
Alvinn: an autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew L. Maas, J. Bagnell, and A. Dey · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell · 2011
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, M. Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes, 2014
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, J. Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Model-based adversarial imitation learning
Nir Baram, Oron Anschel, and Shie Mannor · 2016
Earlier work this paper cites.
Variational inference: A review for statisticians
David M. Blei, A. Kucukelbir, and Jon D. McAuliffe · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, P. Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Infogail: Interpretable imitation learning from visual demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon · 2017
Earlier work this paper cites.
Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets, 2017
Karol Hausman, Yevgen Chebotar, Stefan Schaal, Gaurav Sukhatme, and Joseph Lim · 2017
Earlier work this paper cites.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
Maximilian Karl, Maximilian Soelch, Justin Bayer, and Patrick van der Smagt · 2017
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Earlier work this paper cites.
Deepmind control suite, 2018
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Cited alongside, same era.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I. Jordan, Joseph E. Gonzalez, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Tom Everitt and Marcus Hutter · 2019
Cited alongside, same era.
A Game Theoretic Framework for Model-Based Reinforcement Learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Later among the works it cites.
Learning belief representations for imitation learning in pomdps
Tanmay Gangwani, Joel Lehman, Qiang Liu, and Jian Peng · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X. Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2019
Cited alongside, same era.
Imitation learning as f-divergence minimization
Liyiming Ke, Matt Barnes, W. Sun, Gilwoo Lee, Sanjiban Choudhury, and S. Srinivasa · 2019
Cited alongside, same era.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2019
Cited alongside, same era.
Sample-efficient imitation learning via generative adversarial nets
Lionel Blondé and Alexandros Kalousis · 2019
Cited alongside, same era.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and K. Simonyan · 2019
Cited alongside, same era.
Solar: Deep structured representations for model-based reinforcement learning
Marvin Zhang, Sharad Vikram, Laura Smith, Pieter Abbeel, Matthew J. Johnson, and Sergey Levine · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Cited alongside, same era.
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Lipschitzness is all you need to tame off-policy generative adversarial imitation learning, 2020
Lionel Blondé, Pablo Strasser, and Alexandros Kalousis · 2020
Later among the works it cites.
Model-based inverse reinforcement learning from visual demonstrations
Neha Das, Sarah Bechtle, Todor Davchev, Dinesh Jayaraman, Akshara Rai, and Franziska Meier · 2020
Later among the works it cites.
Morel : Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Offline reinforcement learning from images with latent space models
Rafael Rafailov, Tianhe Yu, Aravind Rajeswaran, and Chelsea Finn · 2020
Later among the works it cites.
Feedback in imitation learning: The three regimes of covariate shift
Jonathan Spencer, Sanjiban Choudhury, Arun Venkatraman, Brian Ziebart, and J. Andrew Bagnell · 2021
Closest in time.
Of moments and matching: Trade-offs and treatments in imitation learning
Gokul Swamy, Sanjiban Choudhury, Zhiwei Steven Wu, and J. Andrew Bagnell · 2021
Closest in time.
Replacing rewards with examples: Example-based policy search via recursive classification, 2021
Benjamin Eysenbach, Sergey Levine, and Ruslan Salakhutdinov · 2021
Closest in time.
Towards learning to imitate from a single video demonstration, 2021
Glen Berseth, Florian Golemo, and Christopher Pal · 2021
Closest in time.
Mobile: Model-based imitation learning from observation alone, 2021
Rahul Kidambi, Jonathan Chang, and Wen Sun · 2021
Closest in time.
Mitigating covariate shift in imitation learning via offline data without great coverage, 2021
Jonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun · 2021
Closest in time.