Fetching the paper…
Reading the bibliography…
The Teacher-Student Framework (TSF) is a reinforcement learning setting where a teacher agent guards the training of a student agent by intervening and providing online demonstrations.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Teaching on a budget: Agents advising agents in reinforcement learning
Lisa Torrey and Matthew Taylor · 2013
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stéphane Ross and J. Andrew Bagnell · 2014
Earlier work this paper cites.
Teacher-student framework: a reinforcement learning approach
Matthieu Zimmer, Paolo Viappiani, and Paul Weng · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Agent-agnostic human-in-the-loop reinforcement learning
David Abel, John Salvatier, Andreas Stuhlmüller, and Owain Evans · 2017
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Earlier work this paper cites.
Collaborative deep reinforcement learning
Kaixiang Lin, Shu Wang, and Jiayu Zhou · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J. Gordon, Byron Boots, and J. Andrew Bagnell · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Policy distillation
Andrei A Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2018
Cited alongside, same era.
A survey on intrinsic motivation in reinforcement learning
Arthur Aubret, Laëtitia Matignon, and Salima Hassas · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Learning to drive by imitation: An overview of deep behavior cloning methods
Abdoulaye O Ly and Moulay Akhloufi · 2020
Later among the works it cites.
Human-in-the-loop imitation learning using remote teleoperation
Ajay Mandlekar, Danfei Xu, Roberto Martín-Martín, Yuke Zhu, Li Fei-Fei, and Silvio Savarese · 2020
Later among the works it cites.
Non-local policy optimization via diversity-regularized collaborative exploration
Zhenghao Peng, Hao Sun, and Bolei Zhou · 2020
Later among the works it cites.
Learning from interventions: Human-robot interaction as both explicit and implicit feedback
Jonathan Spencer, Sanjiban Choudhury, Matthew Barnes, Matthew Schmittle, Mung Chiang, Peter Ramadge, and Siddhartha Srinivasa · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Hg-dagger: Interactive imitation learning with human experts
Michael Kelly, Chelsea Sidrane, Katherine Driggs-Campbell, and Mykel J Kochenderfer · 2019
Cited alongside, same era.
Discorl: Continual reinforcement learning via policy distillation
René Traoré, Hugo Caselles-Dupré, Timothée Lesort, Te Sun, Guanghang Cai, David Filliat, and Natalia Díaz-Rodríguez · 2019
Cited alongside, same era.
On value discrepancy of imitation learning
Tian Xu, Ziniu Li, and Yang Yu · 2019
Cited alongside, same era.
Uncertainty-aware action advising for deep reinforcement learning agents
Felipe Leno da Silva, Pablo Hernandez-Leal, Bilal Kartal, and Matthew Taylor · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Dual policy distillation
Kwei-Herng Lai, Daochen Zha, Yuening Li, and Xia Hu · 2020
Cited alongside, same era.
Later among the works it cites.
Randomized ensembled double q-learning: Learning fast without a model
Xinyue Chen, Che Wang, Zijian Zhou, and Keith W. Ross · 2021
Later among the works it cites.
Correct me if i am wrong: Interactive learning for robotic manipulation
Eugenio Chisari, Tim Welschehold, Joschka Boedecker, Wolfram Burgard, and Abhinav Valada · 2021
Later among the works it cites.
Safe driving via expert guided policy optimization
Zhenghao Peng, Quanyi Li, Chunxiao Liu, and Bolei Zhou · 2021
Later among the works it cites.
Safe reinforcement learning using advantage-based intervention
Nolan Wagener, Byron Boots, and Ching-An Cheng · 2021
Later among the works it cites.
Confidence-aware imitation learning from demonstrations with varying optimality
Songyuan Zhang, Zhangjie Cao, Dorsa Sadigh, and Yanan Sui · 2021
Later among the works it cites.
Robust domain randomised reinforcement learning through peer-to-peer distillation
Chenyang Zhao and Timothy Hospedales · 2021
Later among the works it cites.
Reil: A framework for reinforced intervention-based imitation learning
Rom Parnichkun, Matthew N Dailey, and Atsushi Yamashita · 2022
Later among the works it cites.
Discriminator-weighted offline imitation learning from suboptimal demonstrations
Haoran Xu, Xianyuan Zhan, Honglei Yin, and Huiling Qin · 2022
Later among the works it cites.