Fetching the paper…
Reading the bibliography…
At first sight it may seem straightforward to use recurrent layers in Deep Reinforcement Learning algorithms to enable agents to make use of memory in the setting of partially observable environments.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 1912
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Applications of the morris water maze in the study of learning and memory
D’Hooge, R. and De Deyn, P. P · 2001
Earlier work this paper cites.
Solving deep memory pomdps with recurrent policy gradients
Wierstra, D., Förster, A., Peters, J., and Schmidhuber, J · 2007
Earlier work this paper cites.
Revisiting design choices in proximal policy optimization
Hsu, C. C., Mendler-Dünner, C., and Hardt, M · 2009
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. J. and Stone, P · 2015
Earlier work this paper cites.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Cited alongside, same era.
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castañeda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., Sonnerat, N., Green, T., Deason, L., Leibo, J. Z., Silver, D., Hassabis, D., Kavukcuoglu, K., and Graepel, T · 2018
Cited alongside, same era.
Unity: A general platform for intelligent agents
Juliani, A., Berges, V., Vckay, E., Gao, Y., Henry, H., Mattar, M., and Lange, D · 2018
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T. M., Wu, Y., Powell, G., McGrew, B., and Mordatch, I · 2020
Later among the works it cites.
Implementation matters in deep rl: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Later among the works it cites.
Stale hidden states in ppo-lstm, 2020
Ndousse, K · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
What matters for on-policy deep actor-critic methods? a large-scale study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O · 2021
Later among the works it cites.
Towards mental time travel: a hierarchical memory for reinforcement learning agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D · 2018
Cited alongside, same era.
Generalization of reinforcement learners with working and episodic memory
Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Badia, A. P., Buttimore, G., Deck, C., Leibo, J. Z., and Blundell, C · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., and Munos, R · 2019
Cited alongside, same era.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Cited alongside, same era.
Lampinen, A. K., Chan, S. C. Y., Banino, A., and Hill, F · 2021
Later among the works it cites.
Memory-based deep reinforcement learning for pomdps
Meng, L., Gorbet, R., and Kulic, D · 2021
Later among the works it cites.
Recurrent model-free RL is a strong baseline for many pomdps
Ni, T., Eysenbach, B., and Salakhutdinov, R · 2021
Later among the works it cites.
Robust reinforcement learning on state observations with learned optimal adversary
Zhang, H., Chen, H., Boning, D. S., and Hsieh, C · 2021
Later among the works it cites.