Fetching the paper…
Reading the bibliography…
We address the problem of effectively composing skills to solve sparse-reward tasks in the real world.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
Y. Davidor, Genetic Algorithms and Robotics: A heuristic strategy for optimization . World Scientific, 1991, vol. 1
1991
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, no. 3-4, pp. 229–256, 1992
1992
Earlier work this paper cites.
P. Dayan and G. E. Hinton, “Feudal reinforcement learning,” in Advances in neural information processing systems , 1993, pp. 271–278
1993
Earlier work this paper cites.
L. Chrisman, “Reasoning about probabilistic actions at multiple levels of granularity,” in AAAI Spring Symposium: Decision-Theoretic Planning , 1994
1994
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming . John Wiley & Sons, 1994
1994
Earlier work this paper cites.
E. Gat, “On the role of simulation in the study of autonomous mobile robots,” in AAAI-95 Spring Symposium on Lessons Learned from Implemented Software Architectures for Physical Agents , 1995
1995
Earlier work this paper cites.
M. Asada, S. Noda, S. Tawaratsumida, and K. Hosoda, “Purposive behavior acquisition for a real robot by vision-based reinforcement learning,” Machine learning , vol. 23, no. 2-3, pp. 279–303, 1996
1996
Earlier work this paper cites.
M. Hauskrecht, N. Meuleau, L. P. Kaelbling, T. Dean, and C. Boutilier, “Hierarchical solution of Markov decision processes using macro-actions,” in Proceedings of the Fourteenth conference on Uncertainty in artificial intelligence . Morgan Kaufmann Publishers Inc., 1998, pp. 220–229
1998
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, p. 529, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning , 2015, pp. 1889–1897
2015
Earlier work this paper cites.
I. Mordatch, K. Lowrey, and E. Todorov, “Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids,” in 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2015, pp. 5307–5314
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3431–3440
2015
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2017, pp. 23–30
2017
Later among the works it cites.
F. Sadeghi and S. Levine, “CAD2RL: Real single-image flight without a single real image,” in Robotics: Science and Systems , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 7559–7566
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 1334–1373, 2016
2016
Cited alongside, same era.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in International Conference on Learning Representations , 2016
2016
Cited alongside, same era.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double Q-learning,” in Thirtieth AAAI conference on artificial intelligence , 2016
2016
Cited alongside, same era.
Z. Wang, T. Schaul, M. Hessel, H. Van Hasselt, M. Lanctot, and N. De Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning , 2016
2016
Cited alongside, same era.
W. Masson, P. Ranchod, and G. Konidaris, “Reinforcement learning with parameterized actions,” in Thirtieth AAAI Conference on Artificial Intelligence , 2016
2016
Cited alongside, same era.
M. Hausknecht and P. Stone, “Deep reinforcement learning in parameterized action space,” 2016
2016
Cited alongside, same era.
S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” in 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 2017, pp. 3389–3396
2017
Cited alongside, same era.
J. Andreas, D. Klein, and S. Levine, “Modular multitask reinforcement learning with policy sketches,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2017, pp. 166–175
2017
Cited alongside, same era.
E. Wei, D. Wicke, and S. Luke, “Hierarchical approaches for reinforcement learning in parameterized action space,” in 2018 AAAI Spring Symposium Series , 2018
2018
Later among the works it cites.
K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige et al. , “Using simulation and domain adaptation to improve efficiency of deep robotic grasping,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 4243–4250
2018
Later among the works it cites.
K. Frans, J. Ho, X. Chen, P. Abbeel, and J. Schulman, “Meta learning shared hierarchies,” in International Conference on Learning Representations , 2018
2018
Later among the works it cites.
O. Nachum, S. S. Gu, H. Lee, and S. Levine, “Data-efficient hierarchical reinforcement learning,” in Advances in Neural Information Processing Systems , 2018, pp. 3303–3313
2018
Later among the works it cites.
A. Hill, A. Raffin, M. Ernestus, A. Gleave, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu, “Stable baselines,” https://github.com/hill-a/stable-baselines, 2018
2018
Later among the works it cites.
L. Fan, Y. Zhu, J. Zhu, Z. Liu, O. Zeng, A. Gupta, J. Creus-Costa, S. Savarese, and L. Fei-Fei, “SURREAL: Open-source reinforcement learning framework and robot manipulation benchmark,” in Conference on Robot Learning , 2018
2018
Later among the works it cites.