Fetching the paper…
Reading the bibliography…
The application of reinforcement learning (RL) in robotic control is still limited in the environments with sparse and delayed rewards.
J. A. Bagnell and B. D. Ziebart, “Modeling purposeful adaptive behavior with the principle of maximum causal entropy,” 2010
2010
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pp. 5026–5033, 2012
2012
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, pp. 529–533, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. R. Baker, M. Lai, A. Bolton, Y. Chen, T. P. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, “Mastering the game of go without human knowledge,” Nature , vol. 550, pp. 354–359, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , pp. 488–489, 2017
2017
Earlier work this paper cites.
J. Oh, Y. Guo, S. Singh, and H. Lee, “Self-imitation learning,” in ICML , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter, “Learning agile and dynamic motor skills for legged robots,” Science Robotics , vol. 4, 2019
Y. Guo, J. wook Choi, M. Moczulski, S. Bengio, M. Norouzi, and H. Lee, “Self-imitation learning via trajectory-conditioned policy for hard-exploration tasks,” arXiv: Learning , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Closest in time.
T. Li, K. Srinivasan, M. Q.-H. Meng, W. Yuan, and J. Bohg, “Learning hierarchical control for robust in-hand manipulation,” 2020 IEEE International Conference on Robotics and Automation (ICRA) , pp. 8855–8862, 2020
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar, “Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost,” 2019 International Conference on Robotics and Automation (ICRA) , pp. 3651–3657, 2019
2019
Cited alongside, same era.
J. Choi, K. sik Park, M. Kim, and S. Seok, “Deep reinforcement learning of navigation in a complex and crowded environment with a limited field of view,” 2019 International Conference on Robotics and Automation (ICRA) , pp. 5993–6000, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Reddy, A. D. Dragan, and S. Levine, “Sqil: Imitation learning via reinforcement learning with sparse rewards,” arXiv: Learning , 2020
2020
Closest in time.
A. Singh, E. Jang, A. Irpan, D. Kappler, M. Dalal, S. Levine, M. Khansari, and C. Finn, “Scalable multi-task imitation learning with autonomous improvement,” 2020 IEEE International Conference on Robotics and Automation (ICRA) , pp. 2167–2173, 2020
2020
Closest in time.
K. Brantley, W. Sun, and M. Henaff, “Disagreement-regularized imitation learning,” in ICLR , 2020
2020
Closest in time.