Fetching the paper…
Reading the bibliography…
Conventional reinforcement learning (RL) typically determines an appropriate primitive action at each timestep.
Macro action selection with deep reinforcement learning in StarCraft. In
S. Xu, H. Kuang, Z. Zhi, R. Hu, Y. Liu, and H. Sun. 2019 · 1908
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning. In
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Planning in a hierarchy of abstraction spaces
E. D. Sacerdoti. 1974 · 1974
Earlier work this paper cites.
Macro-operators: A weak method for learning
R. E. Korf. 1985 · 1985
Earlier work this paper cites.
Explanation-based learning: An alternative view
G. DeJong and R. Mooney. 1986 · 1986
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Hierarchical learning in stochastic domains: Preliminary results. In
L. P. Kaelbling. 1993 · 1993
Earlier work this paper cites.
An introduction to genetic algorithms
Melanie Mitchell. 1998 · 1998
Earlier work this paper cites.
Introduction to reinforcement learning . Vol. 135
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Evolutionary algorithms for reinforcement learning
D. E. Moriarty, A. C. Schultz, and J. J. Grefenstette. 1999 · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh. 1999 · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
Macro-FF: Improving AI planning with automatically learned macro-operators
A. Botea, M. Enzenberger, M. Müller, and J. Schaeffer. 2005 · 2005
Cited alongside, same era.
Marvin: A heuristic search planner with online macro-action learning
A. I. Coles and A. J. Smith. 2007 · 2007
Cited alongside, same era.
Learning macro-actions for arbitrary planners and domains. In
M. A. H. Newton, J. Levine, M. Fox, and D. Long. 2007 · 2007
Cited alongside, same era.
Feedback and control systems
Joseph J DiStefano, Allen R Stubberud, and Ivan J Williams. 2012 · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. 2013 · 2013
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction. In
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell. 2017 · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever. 2017 · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. 2017 · 2017
Later among the works it cites.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros. 2018a · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller. 2013 · 2013
Cited alongside, same era.
Solving Large-Scale Planning Problems by Decomposition and Macro Generation. In
M. Asai and A. Fukunaga. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning with macro-actions
I. P. Durugkar, C. Rosenbaum, S. Dernbach, and S. Mahadevan. 2016 · 2016
Cited alongside, same era.
VIME: Variational information maximizing exploration. In
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel. 2016 · 2016
Cited alongside, same era.
Count-based exploration with neural density models. In
Georg Ostrovski, Marc G Bellemare, Aäron Oord, and Rémi Munos. 2017 · 2017
Cited alongside, same era.
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. 2018b · 2018
Later among the works it cites.
Stable baselines
A. Hill, A. Raffin, M. Ernestus, A. Gleave, R. Traore, et al · 2018
Later among the works it cites.
Accelerated methods for deep reinforcement learning
Adam Stooke and Pieter Abbeel. 2018 · 2018
Later among the works it cites.
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. O. Stanley, and J. Clune. 2018 · 2018
Later among the works it cites.
Improving domain-independent planning via critical section macro-operators. In
L. Chrpa and M. Vallati. 2019 · 2019
Closest in time.
Macro action reinforcement learning with sequence disentanglement using variational autoencoder
K. Heecheol, M. Yamada, K. Miyoshi, and H. Yamakawa. 2019 · 2019
Closest in time.
Options of interest: Temporal abstraction with interest functions. In
Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, and Doina Precup. 2020 · 2020
Closest in time.