Fetching the paper…
Reading the bibliography…
Although reinforcement learning methods offer a powerful framework for automatic skill acquisition, for practical learning-based control problems in domains such as robotics, imitation learning often provides a more convenient and accessible alternative.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, and Jan Peters · 1935
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2004
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2004
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D. Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2008
Earlier work this paper cites.
Robot Programming by Demonstration , pp. 1371–1394
Aude Billard, Sylvain Calinon, Rüdiger Dillmann, and Stefan Schaal · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J. Andrew Bagnell · 2014
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Dart: Noise injection for robust imitation learning, 2017
Michael Laskey, Jonathan Lee, Roy Fox, Anca Dragan, and Ken Goldberg · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J. Gordon, Byron Boots, and J. Andrew Bagnell · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Mel Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Earlier work this paper cites.
Fast policy learning through imitation and reinforcement
Ching-An Cheng, Xinyan Yan, Nolan Wagener, and Byron Boots · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Cited alongside, same era.
Hg-dagger: Interactive imitation learning with human experts
Michael Kelly, Chelsea Sidrane, K. Driggs-Campbell, and Mykel J. Kochenderfer · 2018
Cited alongside, same era.
Ensembledagger: A bayesian approach to safe imitation learning
Kunal Menda, K. Driggs-Campbell, and Mykel J. Kochenderfer · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Replacing rewards with examples: Example-based policy search via recursive classification
Ben Eysenbach, Sergey Levine, and Russ R Salakhutdinov · 2021
Later among the works it cites.
Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning
Ryan Hoque, Ashwin Balakrishna, Ellen R. Novoseller, Albert Wilcox, Daniel S. Brown, and Ken Goldberg · 2021
Later among the works it cites.
Jianlan Luo, Oleg O. Sushkov, Rugile Pevceviciute, Wenzhao Lian, Chang Su, Mel Vecerík, Ning Ye, Stefan Schaal, and Jonathan Scholz · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Cited alongside, same era.
Truncated horizon policy search: Combining reinforcement learning and imitation learning
Wen Sun, J. Andrew Bagnell, and Byron Boots · 2018
Cited alongside, same era.
Mo’states mo’problems: Emergency stop mechanisms from observation
Samuel Ainsworth, Matt Barnes, and Siddhartha Srinivasa · 2019
Cited alongside, same era.
Reinforcement learning and optimal control
Dimitri Bertsekas · 2019
Cited alongside, same era.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2019
Cited alongside, same era.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Cited alongside, same era.
Tengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong, and Yu Bai · 2021
Later among the works it cites.
Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity
Abhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham Kakade, and Sergey Levine · 2022
Later among the works it cites.
Fleet-dagger: Interactive robot fleet learning with scalable human supervision
Ryan Hoque, Lawrence Yunliang Chen, Satvik Sharma, K Dharmarajan, Brijen Thananjeyan, P. Abbeel, and Ken Goldberg · 2022
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Seunghyun Lee, Younggyo Seo, Kimin Lee, Pieter Abbeel, and Jinwoo Shin · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Yuda Song, Yifei Zhou, Ayush Sekhari, J Andrew Bagnell, Akshay Krishnamurthy, and Wen Sun · 2022
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Tengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong, and Yu Bai · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
Philip J. Ball, Laura Smith, Ilya Kostrikov, and Sergey Levine · 2023
Closest in time.
Understanding the complexity gains of single-task rl with a curriculum
Qiyang Li, Yuexiang Zhai, Yi Ma, and Sergey Levine · 2023
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Mitsuhiko Nakamoto, Yuexiang Zhai, Anikait Singh, Max Sobol Mark, Yi Ma, Chelsea Finn, Aviral Kumar, and Sergey Levine · 2023
Closest in time.
Guarded policy optimization with imperfect online demonstrations, 2023
Zhenghai Xue, Zhenghao Peng, Quanyi Li, Zhihan Liu, and Bolei Zhou · 2023
Closest in time.