Fetching the paper…
Reading the bibliography…
Learning dexterous manipulation in high-dimensional state-action spaces is an important open challenge with exploration presenting a major bottleneck.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Algorithms for Sequential Decision-Making
M. L. Littman · 1996
Earlier work this paper cites.
Robotics: Modelling, Planning and Control
B. Siciliano, L. Sciavicco, L. Villani, and G. Oriolo · 2008
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mülling, and Y. Altün · 2010
Earlier work this paper cites.
No-regret reductions for imitation learning and structured prediction
S. Ross, G. J. Gordon, and J. A. Bagnell · 2010
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
K. Rawlik, M. Toussaint, and S. Vijayakumar · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
S. Ross and J. A. Bagnell · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Learning to search better than your teacher
K.-W. Chang, A. Krishnamurthy, A. Agarwal, H. Daumé III, and J. Langford · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. A. Riedmiller · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2017
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
Deeply AggreVaTeD: Differentiable imitation learning for sequential prediction
W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell · 2017
Cited alongside, same era.
Learning from demonstrations for real world reinforcement learning
T. Hester, M. Vecerík, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. Agapiou, J. Z. Leibo, and A. Gruslys · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Y. W. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu · 2017
Self-supervised sim-to-real adaptation for visual robotic manipulation, 2019
R. Jeong, Y. Aytar, D. Khosid, Y. Zhou, J. Kay, T. Lampe, K. Bousmalis, and F. Nori · 2019
Later among the works it cites.
Ac-teach: A bayesian actor-critic method for policy learning with an ensemble of suboptimal teachers, 2019
A. Kurenkov, A. Mandlekar, R. Martin-Martin, S. Savarese, and A. Garg · 2019
Later among the works it cites.
Composing entropic policies using divergence correction
J. Hunt, A. Barreto, T. Lillicrap, and N. Heess · 2019
Later among the works it cites.
Information asymmetry in KL-regularized RL
A. Galashov, S. Jayakumar, L. Hasenclever, D. Tirumala, J. Schwarz, G. Desjardins, W. M. Czarnecki, Y. W. Teh, R. Pascanu, and N. Heess · 2019
Later among the works it cites.
Exploiting hierarchy for learning and transfer in kl-regularized rl, 2019
D. Tirumala, H. Noh, A. Galashov, L. Hasenclever, A. Ahuja, G. Wayne, R. Pascanu, Y. W. Teh, and N. Heess · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Revisiting the softmax bellman operator: New benefits and new perspective, 2018
Z. Song, R. E. Parr, and L. Carin · 2018
Cited alongside, same era.
Residual reinforcement learning for robot control, 2018
T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Deepmind control suite, 2018
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller · 2018
Cited alongside, same era.
Sim-to-real reinforcement learning for deformable object manipulation
J. Matas, S. James, and A. J. Davison · 2018
Cited alongside, same era.
A. Novikov, Z. Wang, K. Zolna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, C. Gulcehre, N. Heess, and N. de Freitas · 2020
Closest in time.
Accelerating online reinforcement learning with offline datasets, 2020
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Closest in time.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller · 2020
Closest in time.
Q-learning in enormous action spaces via amortized approximate maximization
T. Van de Wiele, D. Warde-Farley, A. Mnih, and V. Mnih · 2020
Closest in time.
Importance weighted policy learning and adaption, 2020
A. Galashov, J. Sygnowski, G. Desjardins, J. Humplik, L. Hasenclever, R. Jeong, Y. W. Teh, and N. Heess · 2020
Closest in time.
Acme: A research framework for distributed reinforcement learning, 2020
M. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, S. Henderson, A. Novikov, S. G. Colmenarejo, S. Cabi, C. Gulcehre, T. L. Paine, A. Cowie, Z. Wang, B. Piot, and N. de Freitas · 2020
Closest in time.
Soft actor-critic (sac) implementation in pytorch
D. Yarats and I. Kostrikov · 2020
Closest in time.