Fetching the paper…
Reading the bibliography…
Imitation learning (IL) is a popular paradigm for training policies in robotic systems when specifying the reward function is difficult.
Markov Games as a Framework for Multi-Agent Reinforcement Learning. In International Conference on Machine Learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal. 1999 · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation.. In Advances in Neural Information Processing Systems
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning. In International Conference on Machine Learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Approximate equivalence of Markov decision processes
Eyal Even-Dar and Yishay Mansour. 2003 · 2003
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar. 2005 · 2005
Earlier work this paper cites.
Robust reinforcement learning
Jun Morimoto and Kenji Doya. 2005 · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui. 2005 · 2005
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning. In AAAI Conference on Artificial Intelligence
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey. 2008 · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart. 2010 · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control. In IEEE/RSJ International Conference on Intelligent Robots and Systems
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
Feedback control theory
John C Doyle, Bruce A Francis, and Allen R Tannenbaum. 2013 · 2013
Earlier work this paper cites.
Generative Adversarial Networks. In Advances in Neural Information Processing Systems
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms. In International Conference on Machine Learning
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization. In International Conference on Machine Learning
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Avoiding wireheading with value reinforcement learning. In International Conference on Artificial General Intelligence
Tom Everitt and Marcus Hutter. 2016 · 2016
Imitating latent policies from observation. In International Conference on Machine Learning
Ashley Edwards, Himanshu Sahni, Yannick Schroecker, and Charles Isbell. 2019 · 2019
Later among the works it cites.
Action Robust Reinforcement Learning and Applications in Continuous Control. In International Conference on Machine Learning
Chen Tessler, Yonathan Efroni, and Shie Mannor. 2019 · 2019
Later among the works it cites.
Imitation learning from observations by minimizing inverse dynamics disagreement. In Advances in Neural Information Processing Systems
Chao Yang, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Huaping Liu, Junzhou Huang, and Chuang Gan. 2019 · 2019
Later among the works it cites.
An Imitation from Observation Approach to Transfer Learning with Dynamics Mismatch. In Advances in Neural Information Processing Systems
Siddharth Desai, Ishan Durugkar, Haresh Karnan, Garrett Warnell, Josiah Hanna, and Peter Stone. 2020 · 2020
Later among the works it cites.
State-only Imitation with Transition Dynamics Mismatch. In International Conference on Learning Representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative adversarial imitation learning. In Advances in Neural Information Processing Systems
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. 2016 · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning. In International Conference on Learning Representations
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Cited alongside, same era.
Robust Adversarial Reinforcement Learning. In International Conference on Machine Learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta. 2017 · 2017
Cited alongside, same era.
EPOpt: Learning Robust Neural Network Policies Using Model Ensembles. In International Conference on Learning Representations
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Learning Robust Rewards with Adversarial Inverse Reinforcement Learning. In International Conference on Learning Representations
Justin Fu, Katie Luo, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Tanmay Gangwani and Jian Peng. 2020 · 2020
Later among the works it cites.
DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System. In IEEE International Conference on Robotics and Automation
Ankur Handa, Karl Van Wyk, Wei Yang, Jacky Liang, Yu-Wei Chao, Qian Wan, Stan Birchfield, Nathan Ratliff, and Dieter Fox. 2020 · 2020
Later among the works it cites.
Offline Imitation Learning with a Misspecified Simulator. In Advances in Neural Information Processing Systems
Shengyi Jiang, Jingcheng Pang, and Yang Yu. 2020 · 2020
Later among the works it cites.
Robust reinforcement learning via adversarial training with langevin dynamics. In Advances in Neural Information Processing Systems
Parameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland, Cheng Shi, and Volkan Cevher. 2020 · 2020
Later among the works it cites.
State Alignment-based Imitation Learning. In International Conference on Learning Representations
Fangchen Liu, Zhan Ling, Tongzhou Mu, and Hao Su. 2020 · 2020
Later among the works it cites.
Robust Reinforcement Learning for Continuous Control with Model Misspecification. In International Conference on Learning Representations
Daniel J Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki, Jost Tobias Springenberg, Yuanyuan Shi, Jackie Kay, Todd Hester, Timothy Mann, and Martin Riedmiller. 2020 · 2020
Later among the works it cites.
State-only imitation learning for dexterous manipulation
Ilija Radosavovic, Xiaolong Wang, Lerrel Pinto, and Jitendra Malik. 2020 · 2020
Later among the works it cites.
Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey. In IEEE Symposium Series on Computational Intelligence
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund. 2020 · 2020
Later among the works it cites.
Robust Inverse Reinforcement Learning under Transition Dynamics Mismatch. In Advances in Neural Information Processing Systems
Luca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Adrian Weller, and Volkan Cevher. 2021 · 2021
Later among the works it cites.