Fetching the paper…
Reading the bibliography…
Robotic skills can be learned via imitation learning (IL) using user-provided demonstrations, or via reinforcement learning (RL) using large amountsof autonomously collected experience.Both methods have complementarystrengths and weaknesses: RL can reach a high level of performance, but requiresexploration, which can be very time consuming and unsafe; IL does not requireexploration, but only learns skills that are as good as the provided demonstrations.Can a single method combine the strengths of both approaches? A number ofprior methods have aimed to address this question, proposing a variety of tech-niques that integrate elements of IL and RL.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 1910
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
OpenAI, C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 1912
Earlier work this paper cites.
Learning from demonstration
S. Schaal et al · 1997
Earlier work this paper cites.
The Cross Entropy Method: A Unified Approach To Combinatorial Optimization, Monte-Carlo Simulation (Information Science and Statistics)
R. Y. Rubinstein and D. P. Kroese · 2004
Earlier work this paper cites.
Imitation learning for locomotion and manipulation
N. Ratliff, J. A. Bagnell, and S. S. Srinivasa · 2007
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
On learning, representing and generalizing a task in a humanoid robot
S. Calinon, F. Guenter, and A. Billard · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Earlier work this paper cites.
Imitation and reinforcement learning for motor primitives with perceptual coupling
J. Kober, B. Mohler, and J. Peters · 2010
Earlier work this paper cites.
Reinforcement learning of motor skills in high dimensions: A path integral approach
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
Learning stable non-linear dynamical systems with gaussian mixture models
S. Khansari-Zadeh and A. Billard · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Learning to select and generalize striking movements in robot table tennis
K. Mülling, J. Kober, O. Kroemer, and J. Peters · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
R. Laroche, P. Trichelair, and R. T. d. Combes · 2017
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Cited alongside, same era.
Pretraining deep actor-critic reinforcement learning algorithms with expert demonstrations
X. Zhang and H. Ma · 2018
Cited alongside, same era.
Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
G. Kahn, A. Villaflor, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Action image representation: Learning scalable deep grasping policies with zero real world data
M. Khansari, D. Kappler, J. Luo, J. Bingham, and M. Kalakrishnan · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Behavior regularized offline reinforcement learning, 2020
Y. Wu, G. Tucker, and O. Nachum · 2020
Later among the works it cites.
Trajectory-wise control variates for variance reduction in policy gradient methods
C.-A. Cheng, X. Yan, and B. Boots · 2020
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Rl-cyclegan: Improving deep-rl robotics with simulation-to-real
A. Irpan, C. Harris, J. Ibarz, K. Rao, M. Khansari, and S. Levine · 2020
Later among the works it cites.
Long-range indoor navigation with prm-rl
A. Francis, A. Faust, H.-T. L. Chiang, J. Hsu, J. C. Kew, M. Fiser, and T.-W. E. Lee · 2020
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Closest in time.