Fetching the paper…
Reading the bibliography…
Eigenoptions (EOs) have been recently introduced as a promising idea for generating a diverse set of options through the graph Laplacian, having been shown to allow efficient exploration.
Automatic Discovery of Subgoals in Reinforcement Learning using Diverse Density
A. McGovern and A. G. Barto · 2001
Earlier work this paper cites.
Using Relative Novelty to Identify Useful Temporal Abstractions in Reinforcement Learning
Ö. Simsek and A. G. Barto · 2004
Earlier work this paper cites.
Skill Discovery in Continuous Reinforcement Learning Domains using Skill Chaining
G. Konidaris and A. G. Barto · 2009
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Analysis of Watson’s strategies for playing Jeopardy!
G. Tesauro, D. Gondek, J. Lenchner, J. Fan, and J. M. Prager · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Probabilistic Inference for Determining Options in Reinforcement Learning
C. Daniel, H. van Hoof, J. Peters, and G. Neumann · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Adaptive Skills Adaptive Partitions (ASAP)
D. J. Mankowitz, T. A. Mann, and S. Mannor · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
The option-critic architecture
P. Bacon, J. Harb, and D. Precup · 2017
Cited alongside, same era.
Socially aware motion planning with deep reinforcement learning
Decentralized non-communicating multiagent collision avoidance with deep reinforcement learning
Y. Chen, M. Liu, M. Everett, and J. P. How · 2017
Closest in time.
Stochastic Neural Networks for Hierarchical Reinforcement Learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Closest in time.
Reinforcement Learning with Unsupervised Auxiliary Tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
A laplacian framework for option discovery in reinforcement learning
M. C. Machado, M. G. Bellemare, and M. H. Bowling · 2017
Closest in time.
A deep reinforcement learning chatbot
I. V. Serban, C. Sankar, M. Germain, S. Zhang, Z. Lin, S. Subramanian, T. Kim, M. Pieper, S. Chandar, N. R. Ke, S. Mudumba, A. de Brebisson, J. M. R. Sotelo, D. Suhubdy, V. Michalski, A. Nguyen, J. Pineau, and Y. Bengio · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Chen, M. Everett, M. Liu, and J. P. How · 2017
Cited alongside, same era.
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling
Cited in the paper.
Eigenoption Discovery through the Deep Successor Representation
M. C. Machado, C. Rosenbaum, X. Guo, M. Liu, G. Tesauro, and M. Campbell
Cited in the paper.
Proto-value Functions: A Laplacian Framework for Learning Representation and Control in Markov Decision Processes
S. Mahadevan and M. Maggioni
Cited in the paper.
Proto-value functions: A Laplacian framework for learning representation and control in Markov decision processes
S. Mahadevan and M. Maggioni
Cited in the paper.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour
Cited in the paper.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
R. S. Sutton, D. Precup, and S. P. Singh
Cited in the paper.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis · 2017
Closest in time.