Fetching the paper…
Reading the bibliography…
In most practical applications of reinforcement learning, it is untenable to maintain direct estimates for individual states; in continuous-state systems, it is impossible.
The value function polytope in reinforcement learning
Dadashi, R.; Taïga, A. A.; Roux, N. L.; Schuurmans, D.; and Bellemare, M. G. 2019 · 1901
Earlier work this paper cites.
Efficient model-free reinforcement learning in metric spaces
Song, Z.; and Sun, W. 2019 · 1905
Earlier work this paper cites.
Real Analysis
Royden, H. 1968 · 1968
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. 1994 · 1994
Earlier work this paper cites.
On the Generation of Markov Decision Processes
Archibald, T. W.; McKinnon, K. I. M.; and Thomas, L. C. 1995 · 1995
Earlier work this paper cites.
Python reference manual
Van Rossum, G.; and Drake Jr, F. L. 1995 · 1995
Earlier work this paper cites.
Between MDPs and semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Sutton, R.; Precup, D.; and Singh, S. 1999 · 1999
Earlier work this paper cites.
SciPy: Open source scientific tools for Python
Jones, E.; Oliphant, T.; Peterson, P.; et al. 2001 · 2001
Earlier work this paper cites.
Equivalence notions and model minimization in Markov decision processes
Givan, R.; Dean, T.; and Greig, M. 2003 · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Kakade, S.; Kearns, M. J.; and Langford, J. 2003 · 2003
Earlier work this paper cites.
Zooming for Efficient Model-Free Reinforcement Learning in Metric Spaces
Touati, A.; Taiga, A. A.; and Bellemare, M. G. 2020 · 2003
Earlier work this paper cites.
Metrics for finite Markov decision processes
Ferns, N.; Panangaden, P.; and Precup, D. 2004 · 2004
Earlier work this paper cites.
Metrics for Markov Decision Processes with Infinite State Spaces
Ferns, N.; Panangaden, P.; and Precup, D. 2005 · 2005
Earlier work this paper cites.
The Value-Improvement Path: Towards Better Representations for Reinforcement Learning
Dabney, W.; Barreto, A.; Rowland, M.; Dadashi, R.; Quan, J.; Bellemare, M. G.; and Silver, D. 2020 · 2006
Earlier work this paper cites.
Towards a Unified Theory of State Abstraction for MDPs
Li, L.; Walsh, T. J.; and Littman, M. L. 2006 · 2006
Cited alongside, same era.
A guide to NumPy
Oliphant, T. E. 2006 · 2006
Cited alongside, same era.
Learning Invariant Representations for Reinforcement Learning without Reconstruction
Zhang, A.; McAllister, R.; Calandra, R.; Gal, Y.; and Levine, S. 2020 · 2006
Cited alongside, same era.
Matplotlib: A 2D graphics environment
Hunter, J. D. 2007 · 2007
Cited alongside, same era.
Python for scientific computing
Oliphant, T. E. 2007 · 2007
Cited alongside, same era.
Optimal transport: old and new
Villani, C. 2008 · 2008
Cited alongside, same era.
PAC optimal exploration in continuous space Markov decision processes
Pazis, J.; and Parr, R. 2013 · 2013
Later among the works it cites.
Difference of Convex Functions Programming for Reinforcement Learning
Piot, B.; Geist, M.; and Pietquin, O. 2014 · 2014
Later among the works it cites.
Optimal Behavioral Hierarchy
Solway, A.; Diuk, C.; Córdova, N.; Yee, D.; Barto, A. G.; Niv, Y.; and Botvinick, M. M. 2014 · 2014
Later among the works it cites.
MEC—A near-optimal online reinforcement learning algorithm for continuous deterministic systems
Zhao, D.; and Zhu, Y. 2014 · 2014
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
Abadi, M.; Barham, P.; Chen, J.; Chen, Z.; Davis, A.; Dean, J.; Devin, M.; Ghemawat, S.; Irving, G.; Isard, M.; et al. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Notions of state equivalence under partial observability
Castro, P. S.; Panangaden, P.; and Precup, D. 2009 · 2009
Cited alongside, same era.
Introduction to metric and topological spaces
Sutherland, W. A. 2009 · 2009
Cited alongside, same era.
Bounding performance loss in approximate MDP homomorphisms
Taylor, J.; Precup, D.; and Panagaden, P. 2009 · 2009
Cited alongside, same era.
Using bisimulation for policy transfer in MDPs
Castro, P. S.; and Precup, D. 2010 · 2010
Cited alongside, same era.
Continuity and differentiability of expected value functions in dynamic discrete choice models
Norets, A. 2010 · 2010
Cited alongside, same era.
On the locality of action domination in sequential decision making
Rachelson, E.; and Lagoudakis, M. G. 2010 · 2010
Cited alongside, same era.
Abel, D.; Hershkowitz, D. E.; and Littman, M. L. 2017 · 2017
Later among the works it cites.
A Laplacian framework for option discovery in reinforcement learning
Machado, M.; Bellemare, M.; and Bowling, M. 2017 · 2017
Later among the works it cites.
Exploration in structured reinforcement learning
Ok, J.; Proutiere, A.; and Tranos, D. 2018 · 2018
Later among the works it cites.
A Geometric Perspective on Optimal Representations for Reinforcement Learning
Bellemare, M.; Dabney, W.; Dadashi, R.; Ali Taiga, A.; Castro, P. S.; Le Roux, N.; Schuurmans, D.; Lattimore, T.; and Lyle, C. 2019 · 2019
Later among the works it cites.
DeepMDP: Learning Continuous Latent Space Models for Representation Learning
Gelada, C.; Kumar, S.; Buckman, J.; Nachum, O.; and Bellemare, M. G. 2019 · 2019
Later among the works it cites.
Adaptive Discretization for Episodic Reinforcement Learning in Metric Spaces
Sinclair, S. R.; Banerjee, S.; and Yu, C. L. 2019 · 2019
Later among the works it cites.
Scalable methods for computing state similarity in deterministic Markov Decision Processes
Castro, P. S. 2020 · 2020
Later among the works it cites.
Array programming with NumPy
Harris, C. R.; Millman, K. J.; van der Walt, S. J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N. J.; et al. 2020 · 2020
Later among the works it cites.