Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (DRL) has successfully solved various problems recently, typically with a unimodal policy representation.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
Glynn, P. W · 1990
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Reinforcement learning - an introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Mixtures of gaussian processes
Tresp, V · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A. and Barto, A. G · 2001
Earlier work this paper cites.
Multiple model-based reinforcement learning
Doya, K., Samejima, K., Katagiri, K., and Kawato, M · 2002
Earlier work this paper cites.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Menache, I., Mannor, S., and Shimkin, N · 2002
Earlier work this paper cites.
Pattern recognition and machine learning, 5th Edition
Bishop, C. M · 2007
Earlier work this paper cites.
Hybridizing mixtures of experts with support vector machines: Investigation into nonlinear dynamic systems identification
Lima, C. A. M., Coelho, A. L. V., and Zuben, F. J. V · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
Skill characterization based on betweenness
Simsek, Ö. and Barto, A. G · 2008
Earlier work this paper cites.
Viualizing data using t-sne
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Variational mixture of gaussian process experts
Yuan, C. and Neubauer, C · 2008
Earlier work this paper cites.
Hierarchical mixture of classification experts uncovers interactions between brain regions
Yao, B., Walther, D. B., Beck, D. M., and Li, F · 2009
Earlier work this paper cites.
Compositional planning using optimal option models
Silver, D. and Ciosek, K · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Cited alongside, same era.
Auto-encoding variational bayes, 2013
Kingma, D. P. and Welling, M · 2013
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference (extended abstract)
Rawlik, K., Toussaint, M., and Vijayakumar, S · 2013
Cited alongside, same era.
Deep autoregressive networks
Gregor, K., Danihelka, I., Mnih, A., Blundell, C., and Wierstra, D · 2014
Cited alongside, same era.
Neural variational inference and learning in belief networks
Mnih, A. and Gregor, K · 2014
Cited alongside, same era.
Muprop: Unbiased backpropagation for stochastic neural networks
Terrain-adaptive locomotion skills using deep reinforcement learning
Peng, X. B., Berseth, G., and van de Panne, M · 2016
Later among the works it cites.
The option-critic architecture
Bacon, P., Harb, J., and Precup, D · 2017
Later among the works it cites.
The reactor: A sample-efficient actor-critic architecture
Gruslys, A., Azar, M. G., Bellemare, M. G., and Munos, R · 2017
Later among the works it cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T. P., Ghahramani, Z., Turner, R. E., and Levine, S · 2017
Later among the works it cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2017
Later among the works it cites.
Variational mixtures of gaussian processes for classification
Luo, C. and Sun, S · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gu, S., Levine, S., Sutskever, I., and Mnih, A · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., de Freitas, N., and Lanctot, M · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Deep reinforcement learning in parameterized action space
Hausknecht, M. J. and Stone, P · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J · 2016
Cited alongside, same era.
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
REBAR: low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, G., Mnih, A., Maddison, C. J., Lawson, D., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Aurick, Z., Pieter, A., and Sergey, L · 2018
Later among the works it cites.
MCP: learning composable hierarchical control with multiplicative compositional policies
Peng, X. B., Chang, M., Zhang, G., Abbeel, P., and Levine, S · 2019
Later among the works it cites.
DAC: the double actor-critic architecture for learning options
Zhang, S. and Whiteson, S · 2019
Later among the works it cites.
Reinforcement learning from a mixture of interpretable experts
Akrour, R., Tateo, D., and Peters, J · 2020
Later among the works it cites.
Deep Reinforcement Learning
Dong, H., Ding, Z., Zhang, S., and Chang · 2020
Later among the works it cites.
Approximation based variance reduction for reparameterization gradients
Geffner, T. and Domke, J · 2020
Later among the works it cites.