Fetching the paper…
Reading the bibliography…
We examine the problem of learning and planning on high-dimensional domains with long horizons and sparse rewards.
Roles of macro-actions in accelerating reinforcement learning. In Grace Hopper celebration of women in computing
Amy McGovern, Richard S Sutton, and Andrew H Fagg. 1997 · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh. 1999 · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz. 2002 · 2002
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning. In Proceedings of the 25th international conference on Machine learning
Carlos Diuk, Andre Cohen, and Michael L Littman. 2008 · 2008
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. 2013 · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Unifying Count-Based Exploration and Intrinsic Motivation. In NIPS
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos. 2016 · 2016
Cited alongside, same era.
Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation. In NIPS
Tejas D. Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Joshua B. Tenenbaum. 2016 · 2016
Cited alongside, same era.
The Option-Critic Architecture.. In AAAI
Pierre-Luc Bacon, Jean Harb, and Doina Precup. 2017 · 2017
Cited alongside, same era.
Planning with Abstract Markov Decision Processes. In International Conference on Automated Planning and Scheduling
Nakul Gopalan, Marie desJardins, Michael L. Littman, James MacGlashan, Shawn Squire, Stefanie Tellex, John Winder, and Lawson L.S. Wong. 2017 · 2017
Closest in time.
Schema Networks: Zero-shot Transfer with a Generative Causal Model of Intuitive Physics
Ken Kansky, Tom Silver, David A Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, Scott Phoenix, and Dileep George. 2017 · 2017
Closest in time.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Rémi Munos. 2017 · 2017
Closest in time.
FeUdal Networks for Hierarchical Reinforcement Learning. In ICML
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu. 2017 · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep Reinforcement Learning with Double Q-Learning.. In AAAI
Hado Van Hasselt, Arthur Guez, and David Silver. 2016 · 2094
Closest in time.