Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms face two distinct challenges: learning effective representations of past and present observations, and determining how actions influence future returns.
Memory span: a review of the literature
A. B. Blankenship · 1938
Earlier work this paper cites.
Steps toward artificial intelligence
M. Minsky · 1961
Earlier work this paper cites.
The hippocampus as a spatial map: Preliminary evidence from unit activity in the freely-moving rat
J. O’Keefe and J. Dostrovsky · 1971
Earlier work this paper cites.
Memory span: Sources of individual and developmental differences
F. N. Dempster · 1981
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
R. S. Sutton · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Monte carlo pomdps
S. Thrun · 1999
Earlier work this paper cites.
Reinforcement learning with long short-term memory
B. Bakker · 2001
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
K. V. Katsikopoulos and S. E. Engelbrecht · 2003
Earlier work this paper cites.
Modeling the role of working memory and episodic memory in behavioral tasks
E. A. Zilli and M. E. Hasselmo · 2008
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
How to construct deep recurrent neural networks
R. Pascanu, Ç. Gülçehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Difference target propagation
D. Lee, S. Zhang, A. Fischer, and Y. Bengio · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Reinforcement learning and episodic memory in humans and animals: an integrative framework
S. J. Gershman and N. D. Daw · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Earlier work this paper cites.
Optimizing agent behavior over long time scales by transporting value
C. Hung, T. P. Lillicrap, J. Abramson, Y. Wu, M. Mirza, F. Carnevale, A. Ahuja, and G. Wayne · 2018
Earlier work this paper cites.
Sparse attentive backtracking: Temporal credit assignment through reminding
N. R. Ke, A. Goyal, O. Bilaniuk, J. Binas, M. C. Mozer, C. Pal, and Y. Bengio · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
RUDDER: return decomposition for delayed rewards
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter · 2019
Cited alongside, same era.
Amrl: Aggregated memory for reinforcement learning
J. Beck, K. Ciosek, S. Devlin, S. Tschiatschek, C. Zhang, and K. Hofmann · 2019
Cited alongside, same era.
Soft actor-critic for discrete action settings
P. Christodoulou · 2019
Cited alongside, same era.
Generalization of reinforcement learners with working and episodic memory
Synthetic returns for long-term credit assignment
D. Raposo, S. Ritter, A. Santoro, G. Wayne, T. Weber, M. M. Botvinick, H. van Hasselt, and H. F. Song · 2021
Later among the works it cites.
seaborn: statistical data visualization
M. L. Waskom · 2021
Later among the works it cites.
Recurrent off-policy baselines for memory-based continuous control
Z. Yang and H. Nguyen · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
D. Yarats, I. Kostrikov, and R. Fergus · 2021
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Fortunato, M. Tan, R. Faulkner, S. Hansen, A. P. Badia, G. Buttimore, C. Deck, J. Z. Leibo, and C. Blundell · 2019
Cited alongside, same era.
Meta reinforcement learning as task inference
J. Humplik, A. Galashov, L. Hasenclever, P. A. Ortega, Y. W. Teh, and N. Heess · 2019
Cited alongside, same era.
Sequence modeling of temporal credit assignment for episodic reinforcement learning
Y. Liu, Y. Luo, Y. Zhong, X. Chen, Q. Liu, and J. Peng · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew · 2019
Cited alongside, same era.
Self-attentional credit assignment for transfer in reinforcement learning
J. Ferret, R. Marinier, M. Geist, and O. Pietquin · 2020
Cited alongside, same era.
Variational recurrent models for solving partially observable control tasks
D. Han, K. Doya, and J. Tani · 2020
Cited alongside, same era.
C. Chen, Y. Wu, J. Yoon, and S. Ahn · 2022
Later among the works it cites.
Deep transformer q-networks for partially observable reinforcement learning
K. Esslinger, R. Platt, and C. Amato · 2022
Later among the works it cites.
Do as I can, not as I say: Grounding language in robotic affordances
B. Ichter, A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, D. Kalashnikov, S. Levine, Y. Lu, C. Parada, K. Rao, P. Sermanet, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, M. Yan, N. Brown, M. Ahn, O. Cortes, N. Sievers, C. Tan, S. Xu, D. Reyes, J. Rettinghouse, J. Quiambao, P. Pastor, L. Luu, K. Lee, Y. Kuang, S. Jesmonth, N. J. Joshi, K. Jeffrey, R. J. Ruano, J. Hsu, K. Gopalakrishnan, B. David, A. Zeng, and C. K. Fu · 2022
Later among the works it cites.
Multi-game decision transformers
K. Lee, O. Nachum, M. Yang, L. Lee, D. Freeman, S. Guadarrama, I. Fischer, W. Xu, E. Jang, H. Michalewski, and I. Mordatch · 2022
Later among the works it cites.
Transformers are meta-reinforcement learners
L. C. Melo · 2022
Later among the works it cites.
Transformers are sample efficient world models
V. Micheli, E. Alonso, and F. Fleuret · 2022
Later among the works it cites.
Recurrent model-free RL can be a strong baseline for many pomdps
T. Ni, B. Eysenbach, and R. Salakhutdinov · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Evaluating long-term memory in 3d mazes
J. Pasukonis, T. P. Lillicrap, and D. Hafner · 2022
Later among the works it cites.
Learning long-term reward redistribution via randomized return decomposition
Z. Ren, R. Guo, Y. Zhou, and J. Peng · 2022
Later among the works it cites.
Online decision transformer
Q. Zheng, A. Zhang, and A. Grover · 2022
Later among the works it cites.
Transformers in reinforcement learning: A survey
P. Agarwal, A. A. Rahman, P.-L. St-Charles, S. J. Prince, and S. E. Kahou · 2023
Closest in time.
Amago: Scalable in-context reinforcement learning for adaptive agents
J. Grigsby, L. Fan, and Y. Zhu · 2023
Closest in time.
A survey on transformers in reinforcement learning
W. Li, H. Luo, Z. Lin, C. Zhang, Z. Lu, and D. Ye · 2023
Closest in time.
Popgym: Benchmarking partially observable reinforcement learning
S. D. Morad, R. Kortvelesy, M. Bettini, S. Liwicki, and A. Prorok · 2023
Closest in time.
Memory gym: Partially observable challenges to memory-based agents
M. Pleines, M. Pallasch, F. Zimmer, and M. Preuss · 2023
Closest in time.
Transformer-based world models are happy with 100k interactions
J. Robine, M. Höftmann, T. Uelwer, and S. Harmeling · 2023
Closest in time.