Fetching the paper…
Reading the bibliography…
Imitation learning aims to mimic the behavior of experts without explicit reward signals.
A Framework for Behavioural Cloning. In Machine Intelligence 15 . 103–129
Michael Bain and Claude Sammut. 1995 · 1995
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning. In Proceedings of the 17th International Conference on Machine Learning . 663–670
Andrew Y. Ng and Stuart Russell. 2000 · 2000
Earlier work this paper cites.
Asymptopia: An Exposition of Statistical Asymptotic Theory
David Pollard. 2000 · 2000
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning. In Proceedings of the 19th International Conference on Machine Learning . 267–274
Sham M. Kakade and John Langford. 2002 · 2002
Earlier work this paper cites.
Apprenticeship Learning via Inverse Reinforcement Learning. In Proceedings of the 21st International Conference of Machine Learning . 1
Pieter Abbeel and Andrew Y. Ng. 2004 · 2004
Earlier work this paper cites.
Knowledge-Based Kernel Approximation
Olvi L. Mangasarian, Jude W. Shavlik, and Edward W. Wild. 2004 · 2004
Earlier work this paper cites.
Knowledge-based support-vector regression for reinforcement learning
Richard Maclin, Jude Shavlik, Trevor Walker, and Lisa Torrey. 2005 · 2005
Earlier work this paper cites.
Efficient Reductions for Imitation Learning. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics . 661–668
Stéphane Ross and Drew Bagnell. 2010 · 2010
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics , Vol. 15. 627–635
Stéphane Ross, Geoffrey J. Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
Online Learning and Online Convex Optimization
Shai Shalev-Shwartz. 2012 · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. 2013 · 2013
Earlier work this paper cites.
Generative Adversarial Networks. In Advances in Neural Information Processing Systems 27
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Reinforcement and Imitation Learning via Interactive No-Regret Learning
Stéphane Ross and J. Andrew Bagnell. 2014 · 2014
Earlier work this paper cites.
Human-level Control through Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Earlier work this paper cites.
Generative Adversarial Imitation Learning. In Advances in Neural Information Processing Systems 29 . 4565–4573
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
Model-Free Imitation Learning with Policy Optimization. In Proceedings of the 33rd International Conference on Machine Learning . 2760–2769
Jonathan Ho, Jayesh K. Gupta, and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
Training a robot with evaluative feedback and unlabeled guidance signals. In Proceedings of the 25th International Symposium on Robot and Human Interactive Communication . 261–266
Anis Najar, Olivier Sigaud, and Mohamed Chetouani. 2016 · 2016
Cited alongside, same era.
Mastering the Game of Go with Deep Neural Networks and Tree Search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction. In Proceedings of the 34th International Conference on Machine Learning . 3309–3318
Conservative Q-Learning for Offline Reinforcement Learning. In Advances in Neural Information Processing Systems 33
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020 · 2020
Later among the works it cites.
Toward the Fundamental Limits of Imitation Learning. In Advances Neural Information Processing Systems 33
Nived Rajaraman, Lin F. Yang, Jiantao Jiao, and Kannan Ramchandran. 2020 · 2020
Later among the works it cites.
Learning from Interventions: Human-robot Interaction as Both Explicit and Implicit Feedback. In Robotics: Science and Systems XVI
Jonathan C. Spencer, Sanjiban Choudhury, Matt Barnes, Matthew Schmittle, Mung Chiang, Peter J. Ramadge, and Siddhartha S. Srinivasa. 2020 · 2020
Later among the works it cites.
Error Bounds of Imitating Policies and Environments. In Advances in Neural Information Processing Systems 33
Tian Xu, Ziniu Li, and Yang Yu. 2020 · 2020
Later among the works it cites.
IQ-Learn: Inverse Soft-Q Learning for Imitation. In Advances in Neural Information Processing Systems 34 . 4028–4039
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wen Sun, Arun Venkatraman, Geoffrey J. Gordon, Byron Boots, and J. Andrew Bagnell. 2017 · 2017
Cited alongside, same era.
Query-efficient Imitation Learning for End-to-End Autonomous Driving. In Proceedings of the AAAI Conference on Artificial Intelligence 31
Jiakai Zhang and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning . 1856–1865
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
HG-DAgger: Interactive Imitation Learning with Human Experts. In Proceedings of the 35th International Conference on Robotics and Automation . 8077–8083
Michael Kelly, Chelsea Sidrane, Katherine Rose Driggs-Campbell, and Mykel J. Kochenderfer. 2018 · 2018
Cited alongside, same era.
An Overview of Machine Teaching
Xiaojin Zhu, Adish Singla, Sandra Zilles, and Anna N. Rafferty. 2018 · 2018
Cited alongside, same era.
Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In Proceedings of the 36th International Conference on Machine Learning . 1052–1061
Xinshi Chen, Shuang Li, Hui Li, Shaohua Jiang, Yuan Qi, and Le Song. 2019 · 2019
Cited alongside, same era.
Exploring the Limitations of Behavior Cloning for Autonomous Driving. In Proceedings of the International Conference on Computer Vision . 9328–9337
Felipe Codevilla, Eder Santana, Antonio M. López, and Adrien Gaidon. 2019 · 2019
Cited alongside, same era.
When to Trust Your Model: Model-Based Policy Optimization. In Advances in Neural Information Processing Systems 32 . 12498–12509
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. 2019 · 2019
Cited alongside, same era.
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, and Stefano Ermon. 2021 · 2021
Later among the works it cites.
Kimin Lee, Laura M. Smith, and Pieter Abbeel. 2021 · 2021
Later among the works it cites.
Safe Driving via Expert Guided Policy Optimization. In Proceedings of 5th Conference on Robot Learning . 1554–1563
Zhenghao Peng, Quanyi Li, Chunxiao Liu, and Bolei Zhou. 2021 · 2021
Later among the works it cites.
On the Value of Interaction and Function Approximation in Imitation Learning. In Advances in Neural Information Processing Systems 34 . 1325–1336
Nived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu, Jiantao Jiao, and Kannan Ramchandran. 2021 · 2021
Later among the works it cites.
Nearly Minimax Optimal Adversarial Imitation Learning with Known and Unknown Transitions
Tian Xu, Ziniu Li, and Yang Yu. 2021 · 2021
Later among the works it cites.
COMBO: Conservative Offline Model-Based Policy Optimization. In Advances in Neural Information Processing Systems 34 . 28954–28967
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. 2021 · 2021
Later among the works it cites.
Hybrid Value Estimation for Off-policy Evaluation and Offline Reinforcement Learning
Xue-Kun Jin, Xu-Hui Liu, Shengyi Jiang, and Yang Yu. 2022 · 2022
Later among the works it cites.
Metadrive: Composing Diverse Driving Scenarios for Generalizable Reinforcement Learning
Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou. 2022b · 2022
Later among the works it cites.
The Teaching Dimension of Regularized Kernel Learners. In Proceedings of the 39th International Conference on Machine Learning . 17984–18002
Hong Qian, Xu-Hui Liu, Chen-Xi Su, Aimin Zhou, and Yang Yu. 2022 · 2022
Later among the works it cites.
Improve Generated Adversarial Imitation Learning with Reward Variance regularization
Yi-Feng Zhang, Fan-Ming Luo, and Yang Yu. 2022 · 2022
Later among the works it cites.