Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning has been successfully applied to learn robotic control.
L. E. Dubins, “On curves of minimal length with a constraint on average curvature, and with prescribed initial and terminal positions and tangents,” American Journal of mathematics , vol. 79, no. 3, pp. 497–516, 1957
1957
Earlier work this paper cites.
D. A. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,” Neural computation , vol. 3, no. 1, pp. 88–97, 1991
1991
Earlier work this paper cites.
P. Dayan and G. E. Hinton, “Feudal reinforcement learning,” in Advances in Neural Information Processing Systems , S. Hanson, J. Cowan, and C. Giles, Eds., vol. 5. Morgan-Kaufmann, 1992. [Online]. Available: https://proceedings.neurips.cc/paper/1992/file/d14220ee66aeec73c49038385428ec4c-Paper.pdf
1992
Earlier work this paper cites.
L. P. Kaelbling, “Learning to achieve goals,” in IN PROC. OF IJCAI-93 . Morgan Kaufmann, 1993, pp. 1094–1098
1993
Earlier work this paper cites.
T. Blickle and L. Thiele, “A Comparison of Selection Schemes Used in Evolutionary Algorithms,” Evolutionary Computation , vol. 4, no. 4, pp. 361–394, 12 1996. [Online]. Available: https://doi.org/10.1162/evco.1996.4.4.361
1996
Earlier work this paper cites.
S. Russell, “Learning agents for uncertain environments,” in Proceedings of the eleventh annual conference on Computational learning theory , 1998, pp. 101–103
1998
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: an Introduction . MIT press Cambridge, 1998
1998
Earlier work this paper cites.
S. M. LaValle et al. , “Rapidly-exploring random trees: A new tool for path planning,” The annual research report , 1998
1998
Earlier work this paper cites.
A. W. Moore, L. C. Baird, and L. P. Kaelbling, “Multi-value-functions: Efficient automatic action hierarchies for multiple goal mdps,” in IJCAI , 1999
1999
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , vol. 112, no. 1-2, pp. 181–211, 1999
1999
Earlier work this paper cites.
D. Precup, Temporal abstraction in reinforcement learning . University of Massachusetts Amherst, 2000
2000
Earlier work this paper cites.
S. Ross and D. Bagnell, “Efficient reductions for imitation learning,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics , ser. Proceedings of Machine Learning Research, Y. W. Teh and M. Titterington, Eds., vol. 9. Chia Laguna Resort, Sardinia, Italy: PMLR, 13–15 May 2010, pp. 661–668. [Online]. Available: https://proceedings.mlr.press/v9/ross10a.html
2010
Earlier work this paper cites.
K. Y. Levy and N. Shimkin, “Unified inter and intra options learning using policy gradient methods,” in European Workshop on Reinforcement Learning . Springer, 2011, pp. 153–164
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012, pp. 5026–5033
2012
Earlier work this paper cites.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” in International conference on machine learning . PMLR, 2015, pp. 1312–1320
2015
Earlier work this paper cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Vime: Variational information maximizing exploration,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
P.-L. Bacon, J. Harb, and D. Precup, “The option-critic architecture,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 31, no. 1, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
F. Behbahani, K. Shiarlis, X. Chen, V. Kurin, S. Kasewa, C. Stirbu, J. Gomes, S. Paul, F. A. Oliehoek, J. Messias et al. , “Learning from demonstration in the wild,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 775–781
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Levy, G. D. Konidaris, R. W. Platt, and K. Saenko, “Learning multi-level hierarchies with hindsight,” in ICLR , 2019
2019
Later among the works it cites.
B. Eysenbach, R. R. Salakhutdinov, and S. Levine, “Search on the replay buffer: Bridging planning and reinforcement learning,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
X. B. Peng, P. Abbeel, S. Levine, and M. Van de Panne, “Deepmimic: Example-guided deep reinforcement learning of physics-based character skills,” ACM Transactions On Graphics (TOG) , vol. 37, no. 4, pp. 1–14, 2018
2018
Cited alongside, same era.
O. Nachum, S. S. Gu, H. Lee, and S. Levine, “Data-efficient hierarchical reinforcement learning,” Advances in neural information processing systems , vol. 31, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning . PMLR, 2018, pp. 1861–1870
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Bagaria and G. Konidaris, “Option discovery using deep skill chaining,” in International Conference on Learning Representations , 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
G. Matheron, N. Perrin, and O. Sigaud, “Pbcs: Efficient exploration and exploitation using a synergy between reinforcement learning and motion planning,” in International Conference on Artificial Neural Networks . Springer, 2020, pp. 295–307
2020
Later among the works it cites.
A. Kuznetsov, P. Shvechikov, A. Grishin, and D. Vetrov, “Controlling overestimation bias with truncated mixture of continuous distributional quantile critics,” in International Conference on Machine Learning . PMLR, 2020, pp. 5556–5566
2020
Later among the works it cites.
E. Johns, “Coarse-to-fine imitation learning: Robot manipulation from a single demonstration,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 4613–4619
2021
Later among the works it cites.
E. Chane-Sane, C. Schmid, and I. Laptev, “Goal-conditioned reinforcement learning with imagined subgoals,” in International Conference on Machine Learning . PMLR, 2021, pp. 1430–1440
2021
Later among the works it cites.
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune, “First return, then explore,” Nature , vol. 590, no. 7847, pp. 580–586, 2021
2021
Later among the works it cites.
A. Bagaria, J. Senthil, M. Slivinski, and G. Konidaris, “Robustly learning composable options in deep reinforcement learning,” in Proceedings of the 30th International Joint Conference on Artificial Intelligence , 2021
2021
Later among the works it cites.
A. Gupta, J. Yu, T. Z. Zhao, V. Kumar, A. Rovinsky, K. Xu, T. Devlin, and S. Levine, “Reset-free reinforcement learning via multi-task learning: Learning dexterous manipulation behaviors without human intervention,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 6664–6671
2021
Later among the works it cites.
2022
Closest in time.
M. Hutsebaut-Buysse, K. Mets, and S. Latré, “Hierarchical reinforcement learning: A survey and open research challenges,” Machine Learning and Knowledge Extraction , vol. 4, no. 1, pp. 172–221, 2022
2022
Closest in time.
M. M. Contributors, “MuJoCo Menagerie: A collection of high-quality simulation models for MuJoCo,” 2022. [Online]. Available: http://github.com/deepmind/mujoco_menagerie
2022
Closest in time.
N. Perrin-Gilbert, “xpag: a modular reinforcement learning library with jax agents,” 2022. [Online]. Available: https://github.com/perrin-isir/xpag
2022
Closest in time.