Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), state representations are key to dealing with large or continuous state spaces.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H. (2019) · 1902
Earlier work this paper cites.
Analysis of a complex of statistical variables into principal components
Hotelling, H. (1933) · 1933
Earlier work this paper cites.
Dynamic programming princeton university press princeton
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Efficient memory-based learning for robot control
Moore, A. W. (1990) · 1990
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Temporal-difference methods and markov models
Barnard, E. (1993) · 1993
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L. (1995) · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. and Van Roy, B. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
SciPy: open source scientific tools for Python
Jones, E., Oliphant, T., and Peterson, P. (2001) · 2001
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Boyan, J. A. (2002) · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Invariant subspaces of matrices with applications
Gohberg, I., Lancaster, P., and Rodman, L. (2006) · 2006
Earlier work this paper cites.
A guide to NumPy
Oliphant, T. E. (2006) · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D. (2007) · 2007
Earlier work this paper cites.
Python for scientific computing
Oliphant, T. E. (2007) · 2007
Earlier work this paper cites.
Basis function adaptation methods for cost approximation in mdp
Yu, H. and Bertsekas, D. P. (2009) · 2009
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Cited alongside, same era.
Should one compute the temporal difference fix point or minimize the bellman residual? the unified oblique projection view
Scherrer, B. (2010) · 2010
Cited alongside, same era.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Halko, N., Martinsson, P.-G., and Tropp, J. A. (2011) · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
A geometric perspective on optimal representations for reinforcement learning
Bellemare, M., Dabney, W., Dadashi, R., Ali Taiga, A., Castro, P. S., Le Roux, N., Schuurmans, D., Lattimore, T., and Lyle, C. (2019) · 2019
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G. (2019) · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W. (2019) · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The numpy array: a structure for efficient numerical computation
Walt, S. v. d., Colbert, S. C., and Varoquaux, G. (2011) · 2011
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Dann, C., Neumann, G., Peters, J., et al. (2014) · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., and Xiaoqiang, Z. (2016) · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. (2016) · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M. (2016) · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D. (2017) · 2017
Cited alongside, same era.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Later among the works it cites.
Exponentially convergent stochastic k-pca without variance reduction
Tang, C. (2019) · 2019
Later among the works it cites.
Representations for stable off-policy reinforcement learning
Ghosh, D. and Bellemare, M. G. (2020) · 2020
Later among the works it cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Guo, Z. D., Pires, B. A., Piot, B., Grill, J.-B., Altché, F., Munos, R., and Azar, M. G. (2020) · 2020
Later among the works it cites.
Array programming with numpy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., et al. (2020) · 2020
Later among the works it cites.
Gamma-models: Generative temporal difference learning for infinite-horizon prediction
Janner, M., Mordatch, I., and Levine, S. (2020) · 2020
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint
Blier, L., Tallec, C., and Ollivier, Y. (2021) · 2021
Later among the works it cites.
The value-improvement path: Towards better representations for reinforcement learning
Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D. (2021) · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W. (2021) · 2021
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P. (2021) · 2021
Later among the works it cites.
Learning one representation to optimize all rewards
Touati, A. and Ollivier, Y. (2021) · 2021
Later among the works it cites.
Breaking the deadly triad with a target network
Zhang, S., Yao, H., and Whiteson, S. (2021) · 2021
Later among the works it cites.
On the generalization of representations in reinforcement learning
Le Lan, C., Tu, S., Oberman, A., Agarwal, R., and Bellemare, M. G. (2022) · 2022
Later among the works it cites.
Generalised policy improvement with geometric policy composition
Thakoor, S., Rowland, M., Borsa, D., Dabney, W., Munos, R., and Barreto, A. (2022) · 2022
Later among the works it cites.
Proto-value networks: Scaling representation learning with auxiliary tasks
Farebrother, J., Greaves, J., Agarwal, R., Lan, C. L., Goroshin, R., Castro, P. S., and Bellemare, M. G. (2023) · 2023
Closest in time.
A novel stochastic gradient descent algorithm for learning principal subspaces
Le Lan, C., Greaves, J., Farebrother, J., Rowland, M., Pedregosa, F., Agarwal, R., and Bellemare, M. G. (2023) · 2023
Closest in time.
Understanding self-predictive learning for reinforcement learning
Tang, Y., Guo, Z. D., Richemond, P. H., Pires, B. Á., Chandak, Y., Munos, R., Rowland, M., Azar, M. G., Lan, C. L., Lyle, C., et al. (2023) · 2023
Closest in time.