Fetching the paper…
Reading the bibliography…
Real-world tasks are often highly structured.
Using expectation-maximization for reinforcement learning
P. Dayan and G. Hinton · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
T. G Dietterich · 2000
Earlier work this paper cites.
Visualizing data using t-sne
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Fitted q-iteration by advantage weighted regression
G. Neumann and J. Peters · 2009
Earlier work this paper cites.
Discriminative clustering by regularized information maximization
R. Gomes, A. Krause, and P. Perona · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
B. Ziebart · 2010
Earlier work this paper cites.
Policy search for motor primitives in robotics
J. Kober and J. Peters · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, , and Yuval Tassa · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu1, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Openai gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, O. Kroemer, and J. Peters · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
Variational intrinsic control
K. Gregor, D. Rezende, and D. Wierstra · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Learning discrete representations via information maximizing self augmented training
W. Hu, T. Miyato, S. Tokui, E. Matsumoto, and M. Sugiyama · 2017
Later among the works it cites.
InfoGAIL: Interpretable imitation learning from visual demonstrations
Y. Li, J. Song, and S. Ermon · 2017
Later among the works it cites.
FeUdal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Diversity is all you need: Learning diverse skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and Sergey Levine · 2018
Later among the works it cites.
Meta learning shared hierarchies
K. Frans, J. Ho, X. Chen, P. Abbeel, and J. Schulman · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Distributional smoothing with virtual adversarial training
T. Miyato, S. Maeda, M. Koyama, K. Nakae, and S. Ishii · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
The option-critic architecture
P. L. Bacon, J. Harb, and D. Precup · 2017
Cited alongside, same era.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Cited alongside, same era.
Later among the works it cites.
Divide-and-conquer reinforcement learning
D. Ghosh, A. Singh, A. Rajeswaran, V. Kumar, and S. Levine · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Later among the works it cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Later among the works it cites.
Hierarchical policy search via return-weighted density estimation
T. Osa and M. Sugiyama · 2018
Later among the works it cites.
An inference-based policy gradient method for learning options
M. J. A. Smith, H. Van Hoof, and J. Pineau · 2018
Later among the works it cites.