Fetching the paper…
Reading the bibliography…
This paper introduces a set of formally defined and transparent problems for reinforcement learning algorithms with the following characteristics: (1) variable degrees of observability (non-Markov observations), (2) distal and sparse rewards, (3) variable and hierarchical reward structure, (4) multiple-task generation, (5) variable problem complexity.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J. (2019) · 1901
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Mazes, maps, and memory
Olton, D. S. (1979) · 1979
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Sequence learning
Clegg, B. A., DiGirolamo, G. J., and Keele, S. W. (1998) · 1998
Earlier work this paper cites.
Assessment of spatial memory using the t maze
Wenk, G. L. (1998) · 1998
Earlier work this paper cites.
Monte carlo POMDPs
Thrun, S. (2000) · 2000
Earlier work this paper cites.
Solving the distal reward problem through linkage of stdp and dopamine signalinig
Izhikevich, E. M. (2007) · 2007
Earlier work this paper cites.
Lifelong learning of compositional structures
Mendez, J. A. and Eaton, E. (2020) · 2007
Earlier work this paper cites.
Evolutionary advantages of neuromodulated plasticity in dynamic, reward-based scenarios
Soltoggio, A., Bullinaria, J. A., Mattiussi, C., Dürr, P., and Floreano, D. (2008) · 2008
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Earlier work this paper cites.
Rare neural correlations implement robotic conditioning with delayed rewards and disturbances
Soltoggio, A., Lemme, A., Reinhart, F., and Steil, J. J. (2013) · 2013
Earlier work this paper cites.
Solving the distal reward problem with rare correlations
Soltoggio, A. and Steil, J. J. (2013) · 2013
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
Hausknecht, M. and Stone, P. (2015) · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
R L 2 {RL}^{2} : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Amrl: aggregated memory for reinforcement learning
Beck, J., Ciosek, K., Devlin, S., Tschiatschek, S., Zhang, C., and Hofmann, K. (2019) · 2019
Later among the works it cites.
Deep learning for video game playing
Justesen, N., Bontrager, P., Togelius, J., and Risi, S. (2019) · 2019
Later among the works it cites.
Curiosity-bottleneck: Exploration by distilling task-specific novelty
Kim, Y., Nam, W., Kim, H., Kim, J.-H., and Kim, G. (2019) · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. (2019) · 2019
Later among the works it cites.
Fast context adaptation via meta-learning
Zintgraf, L., Shiarli, K., Kurin, V., Hofmann, K., and Whiteson, S. (2019) · 2019
Later among the works it cites.
Evolving inborn knowledge for fast adaptation in dynamic pomdp problems
Ben-Iwhiwhu, E., Ladosz, P., Dick, J., Chen, W.-H., Pilly, P., and Soltoggio, A. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning: An overview
Li, Y. (2017) · 2017
Cited alongside, same era.
Neural map: Structured memory for deep reinforcement learning
Parisotto, E. and Salakhutdinov, R. (2017) · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O. (2018) · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S. (2018) · 2018
Cited alongside, same era.
Differentiable plasticity: training plastic neural networks with backpropagation
Miconi, T., Stanley, K., and Clune, J. (2018) · 2018
Cited alongside, same era.
Born to learn: the inspiration, progress, and future of evolved plastic artificial neural networks
Soltoggio, A., Stanley, K. O., and Risi, S. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Detecting changes and avoiding catastrophic forgetting in dynamic partially observable environments
Dick, J., Ladosz, P., Ben-Iwhiwhu, E., Shimadzu, H., Kinnell, P., Pilly, P. K., Kolouri, S., and Soltoggio, A. (2020) · 2020
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Küttler, H., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Later among the works it cites.
Exploration in deep reinforcement learning: A survey
Ladosz, P., Weng, L., Kim, M., and Oh, H. (2022) · 2022
Later among the works it cites.
A domain-agnostic approach for characterization of lifelong learning systems
Baker, M. M., New, A., Aguilar-Simon, M., Al-Halah, Z., Arnold, S. M., Ben-Iwhiwhu, E., Brna, A. P., Brooks, E., Brown, R. C., Daniels, Z., Daram, A., Delattre, F., Dellana, R., Eaton, E., Fu, H., Grauman, K., Hostetler, J., Iqbal, S., Kent, D., Ketz, N., Kolouri, S., Konidaris, G., Kudithipudi, D., Learned-Miller, E., Lee, S., Littman, M. L., Madireddy, S., Mendez, J. A., Nguyen, E. Q., Piatko, C., Pilly, P. K., Raghavan, A., Rahman, A., Ramakrishnan, S. K., Ratzlaff, N., Soltoggio, A., Stone, P., Sur, I., Tang, Z., Tiwari, S., Vedder, K., Wang, F., Xu, Z., Yanguas-Gil, A., Yedidsion, H., Yu, S., and Vallabha, G. K. (2023) · 2023
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Closest in time.
Deep reinforcement learning with modulated hebbian plus q-network architecture
Ladosz, P., Ben-Iwhiwhu, E., Dick, J., Ketz, N., Kolouri, S., Krichmar, J. L., Pilly, P. K., and Soltoggio, A. (2021) · 2056
Closest in time.