Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) for reinforcement learning (RL) have shown distinct advantages, e.g., solving memory-dependent tasks and meta-learning.
On the theory of the brownian motion
Uhlenbeck, G. E. and Ornstein, L. S · 1930
Earlier work this paper cites.
A Markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Optimal control of Markov processes with incomplete state information
Åström, K. J · 1965
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
Schmidhuber, J · 1991
Earlier work this paper cites.
The highly irregular firing of cortical cells is inconsistent with temporal integration of random epsps
Softky, W. R. and Koch, C · 1993
Earlier work this paper cites.
Learning one more thing
Thrun, S. and Mitchell, T. M · 1994
Earlier work this paper cites.
Identification of a forebrain motor programming network for the learned song of zebra finches
Vu, E. T., Mazurek, M. E., and Kuo, Y.-C · 1994
Earlier work this paper cites.
On bias, variance, 0/1–loss, and the curse-of-dimensionality
Friedman, J. H · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Learning to learn: Introduction and overview
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K · 2000
Earlier work this paper cites.
Time scales in motor learning and development
Newell, K. M., Liu, Y.-T., and Mayer-Kress, G · 2001
Earlier work this paper cites.
Multiple time scales and multiform dynamics in learning to juggle
Huys, R., Daffertshofer, A., and Beek, P. J · 2004
Earlier work this paper cites.
Interacting adaptive processes with different timescales underlie short-term motor learning
Smith, M. A., Ghazizadeh, A., and Shadmehr, R · 2006
Earlier work this paper cites.
Solving deep memory pomdps with recurrent policy gradients
Wierstra, D., Foerster, A., Peters, J., and Schmidhuber, J · 2007
Earlier work this paper cites.
Probabilistic population codes for Bayesian decision making
Beck, J. M., Ma, W. J., Kiani, R., Hanks, T., Churchland, A. K., Roitman, J., Shadlen, M. N., Latham, P. E., and Pouget, A · 2008
Earlier work this paper cites.
Contextual behaviors and internal representations acquired by reinforcement learning with a recurrent neural network in a continuous state and action space task
Utsunomiya, H. and Shibata, K · 2008
Earlier work this paper cites.
Emergence of functional hierarchy in a multiple timescale neural network model: a humanoid robot experiment
Yamashita, Y. and Tani, J · 2008
Earlier work this paper cites.
Frontal cortex and the discovery of abstract action rules
Badre, D., Kayser, A. S., and D’Esposito, M · 2010
Earlier work this paper cites.
Mechanisms of hierarchical reinforcement learning in cortico–striatal circuits 2: Evidence from fmri
Badre, D. and Frank, M. J · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Dopamine neurons learn to encode the long-term value of multiple future rewards
Enomoto, K., Matsumoto, N., Nakai, S., Satoh, T., Sato, T. K., Ueda, Y., Inokawa, H., Haruno, M., and Kimura, M · 2011
Earlier work this paper cites.
Not noisy, just wrong: the role of suboptimal inference in behavioral variability
Beck, J. M., Ma, W. J., Pitkow, X., Latham, P. E., and Pouget, A · 2012
Cited alongside, same era.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S · 2012
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Cited alongside, same era.
Creating a false memory in the hippocampus
Ramirez, S., Liu, X., Lin, P.-A., Suh, J., Pignatelli, M., Redondo, R. L., Ryan, T. J., and Tonegawa, S · 2013
Cited alongside, same era.
Lifelong machine learning systems: Beyond learning algorithms
Silver, D. L., Yang, Q., and Li, L · 2013
Cited alongside, same era.
PAC-inspired option discovery in lifelong reinforcement learning
Brunskill, E. and Li, L · 2014
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2017
Later among the works it cites.
Multi-level discovery of deep options
Fox, R., Krishnan, S., Stoica, I., and Goldberg, K · 2017
Later among the works it cites.
The role of forelimb motor cortex areas in goal directed action in mice
Morandell, K. and Huber, D · 2017
Later among the works it cites.
Distinct timescales of population coding across cortex
Runyan, C. A., Piasini, E., Panzeri, S., and Harvey, C. D · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A hierarchy of intrinsic timescales across primate cortex
Murray, J. D., Bernacchia, A., Freedman, D. J., Romo, R., Wallis, J. D., Cai, X., Padoa-Schioppa, C., Pasternak, T., Seo, H., Lee, D., et al · 2014
Cited alongside, same era.
A large-scale circuit mechanism for hierarchical dynamical processing in the primate cortex
Chaudhuri, R., Knoblauch, K., Gariel, M.-A., Kennedy, H., and Wang, X.-J · 2015
Cited alongside, same era.
A recurrent latent variable model for sequential data
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A. C., and Bengio, Y · 2015
Cited alongside, same era.
Where’s the noise? key features of spontaneous activity and neural variability arise through learning in a deterministic network
Hartmann, C., Lazar, A., Nessler, B., and Triesch, J · 2015
Cited alongside, same era.
Deep recurrent Q-learning for partially observable MDPs
Hausknecht, M. and Stone, P · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D · 2015
Cited alongside, same era.
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2017
Later among the works it cites.
Policy and value transfer in lifelong reinforcement learning
Abel, D., Jinnai, Y., Guo, S. Y., Konidaris, G., and Littman, M · 2018
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P · 2018
Later among the works it cites.
Probabilistic model-agnostic meta-learning
Finn, C., Xu, K., and Levine, S · 2018
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al · 2018
Later among the works it cites.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., and Munos, R · 2018
Later among the works it cites.
Learning abstract options
Riemer, M., Liu, M., and Tesauro, G · 2018
Later among the works it cites.
Prefrontal cortex as a meta-reinforcement learning system
Wang, J. X., Kurth-Nelson, Z., Kumaran, D., Tirumala, D., Soyer, H., Leibo, J. Z., Hassabis, D., and Botvinick, M · 2018
Later among the works it cites.
Bayesian model-agnostic meta-learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S · 2018
Later among the works it cites.
A novel predictive-coding-inspired variational rnn model for online prediction and recognition
Ahmadi, A. and Tani, J · 2019
Closest in time.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Closest in time.
Model-based reinforcement learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Closest in time.