Fetching the paper…
Reading the bibliography…
Learning a good representation is a crucial challenge for Reinforcement Learning (RL) agents.
Predictive representations of state
M. Littman and R. S. Sutton · 2001
Earlier work this paper cites.
Chapter 51 - Alexandr Mikhailovich Lyapunov, Thesis on the stability of motion (1892)
J. Mawhin · 2005
Earlier work this paper cites.
Proto-transfer learning in Markov decision processes using spectral methods
K. Ferguson and S. Mahadevan · 2006
Earlier work this paper cites.
Guide to numpy , volume 1
T. E. Oliphant et al · 2006
Earlier work this paper cites.
The matplotlib user’s guide
J. Hunter and D. Dale · 2007
Earlier work this paper cites.
Ordinary differential equations and dynamical systems
G. Teschl · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Earlier work this paper cites.
OpenAI gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Learning awareness models
B. Amos, L. Dinh, S. Cabi, T. Rothörl, S. G. Colmenarejo, A. Muldal, T. Erez, Y. Tassa, N. de Freitas, and M. Denil · 2018
Earlier work this paper cites.
Eigenoption discovery through the deep successor representation
M. C. Machado, C. Rosenbaum, X. Guo, M. Liu, G. Tesauro, and M. Campbell · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Fast feature selection for linear value function approximation
B. Behzadian, S. Gharatappeh, and M. Petrik · 2019
Cited alongside, same era.
The DeepMind JAX Ecosystem, 2020
DeepMind, I. Babuschkin, K. Baumli, A. Bell, S. Bhupatiraju, J. Bruce, P. Buchlovsky, D. Budden, T. Cai, A. Clark, I. Danihelka, A. Dedieu, C. Fantacci, J. Godwin, C. Jones, R. Hemsley, T. Hennigan, M. Hessel, S. Hou, S. Kapturowski, T. Keck, I. Kemaev, M. King, M. Kunesch, L. Martens, H. Merzic, V. Mikulik, T. Norman, G. Papamakarios, J. Quan, R. Ring, F. Ruiz, A. Sanchez, L. Sartran, R. Schneider, E. Sezener, S. Spencer, S. Srinivasan, M. Stanojević, W. Stokowiec, L. Wang, G. Zhou, and F. Viola · 2020
Cited alongside, same era.
Bootstrap your own latent: A new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko · 2020
Cited alongside, same era.
Bootstrap latent-predictive representations for multitask reinforcement learning
Z. D. Guo, B. A. Pires, B. Piot, J.-B. Grill, F. Altché, R. Munos, and M. G. Azar · 2020
On the effect of auxiliary tasks on representation dynamics
C. Lyle, M. Rowland, G. Ostrovski, and W. Dabney · 2021
Later among the works it cites.
BYOL-Explore: Exploration by bootstrapped prediction
Z. Guo, S. Thakoor, M. Pîslar, B. Avila Pires, F. Altché, C. Tallec, A. Saade, D. Calandriello, J.-B. Grill, Y. Tang, M. Valko, R. Munos, M. Gheshlaghi Azar, and B. Piot · 2022
Later among the works it cites.
gymnax: A JAX-based reinforcement learning environment library, 2022
R. T. Lange · 2022
Later among the works it cites.
Representations and exploration for deep reinforcement learning using singular value decomposition
Y. Chandak, S. Thakoor, Z. D. Guo, Y. Tang, R. Munos, W. Dabney, and D. L. Borsa · 2023
Later among the works it cites.
Minigrid & Miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
M. Chevalier-Boisvert, B. Dai, M. Towers, R. de Lazcano, L. Willems, S. Lahlou, S. Pal, P. S. Castro, and J. Terry · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering Atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Cited alongside, same era.
V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, et al · 2020
Cited alongside, same era.
Scipy 1.0: fundamental algorithms for scientific computing in python
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, et al · 2020
Cited alongside, same era.
Exploring simple Siamese representation learning
X. Chen and K. He · 2021
Cited alongside, same era.
Predicting what you already know helps: Provable self-supervised learning
J. D. Lee, Q. Lei, N. Saunshi, and J. Zhuo · 2021
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2023
J. Heek, A. Levskaya, A. Oliver, M. Ritter, B. Rondepierre, A. Steiner, and M. van Zee · 2023
Later among the works it cites.
Bootstrapped representations in reinforcement learning
C. L. Lan, S. Tu, M. Rowland, A. Harutyunyan, R. Agarwal, M. G. Bellemare, and W. Dabney · 2023
Later among the works it cites.
Spectral decomposition representation for reinforcement learning
T. Ren, T. Zhang, L. Lee, J. E. Gonzalez, D. Schuurmans, and B. Dai · 2023
Later among the works it cites.
Understanding self-predictive learning for reinforcement learning
Y. Tang, Z. D. Guo, P. H. Richemond, B. Á. Pires, Y. Chandak, R. Munos, M. Rowland, M. G. Azar, C. L. Lan, C. Lyle, et al · 2023
Later among the works it cites.
Flashbax: Streamlining experience replay buffers for reinforcement learning with jax, 2023
E. Toledo, L. Midgley, D. Byrne, C. R. Tilbury, M. Macfarlane, C. Courtot, and A. Laterre · 2023
Later among the works it cites.
Bridging state and history representations: Understanding self-predictive RL
T. Ni, B. Eysenbach, E. Seyedsalehi, M. Ma, C. Gehring, A. Mahajan, and P.-L. Bacon · 2024
Closest in time.