Fetching the paper…
Reading the bibliography…
In reinforcement learning, state representations are used to tractably deal with large problem spaces.
Finite continuous time markov chains
John G Kemeny and J Laurie Snell · 1961
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Python reference manual
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
The nature of statistical learning
Vladimir N Vapnik · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S Sutton · 1996
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R.S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Scipy: Open source scientific tools for python
Eric Jones, Travis Oliphant, Pearu Peterson, et al · 2001
Earlier work this paper cites.
Sparse distributed memories for on-line value-based reinforcement learning
Bohdana Ratitch and Doina Precup · 2004
Earlier work this paper cites.
Toeplitz and circulant matrices: A review
Robert M. Gray · 2006
Earlier work this paper cites.
A guide to NumPy , volume 1
Travis E Oliphant · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
John D Hunter · 2007
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Earlier work this paper cites.
Python for scientific computing
Travis E Oliphant · 2007
Earlier work this paper cites.
An analysis of laplacian methods for value function approximation in mdps
Marek Petrik · 2007
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L Littman · 2008
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Compressed least-squares regression
Odalric-Ambrym Maillard and Rémi Munos · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E. Taylor and Peter Stone · 2009
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Cited alongside, same era.
Value function approximation in reinforcement learning using the fourier basis
George D. Konidaris, Sarah Osentoski, and Philip S. Thomas · 2011
Cited alongside, same era.
The numpy array: a structure for efficient numerical computation
Stéfan van der Walt, S Chris Colbert, and Gael Varoquaux · 2011
Cited alongside, same era.
Introduction to probability
Charles Miller Grinstead and James Laurie Snell · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Two-timescale networks for nonlinear value function approximation
Wesley Chung, Somjit Nath, Ajin Joseph, and Martha White · 2018
Later among the works it cites.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
A geometric perspective on optimal representations for reinforcement learning
Marc Bellemare, Will Dabney, Robert Dadashi, Adrien Ali Taiga, Pablo Samuel Castro, Nicolas Le Roux, Dale Schuurmans, Tor Lattimore, and Clare Lyle · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Markov chains: Gibbs fields, Monte Carlo simulation, and queues , volume 31
Pierre Brémaud · 2013
Cited alongside, same era.
Optimal behavioral hierarchy
Alec Solway, Carlos Diuk, Natalia Córdova, Debbie Yee, Andrew G Barto, Yael Niv, and Matthew M Botvinick · 2014
Cited alongside, same era.
Design principles of the hippocampal cognitive map
Kimberly L. Stachenfeld, Matthew Botvinick, and Samuel J. Gershman · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A. Tropp · 2015
Cited alongside, same era.
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Later among the works it cites.
Representation learning on graphs: A reinforcement learning application
Sephora Madjiheurem and Laura Toni · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Later among the works it cites.
The value-improvement path: Towards better representations for reinforcement learning
Will Dabney, André Barreto, Mark Rowland, Robert Dadashi, John Quan, Marc G Bellemare, and David Silver · 2020
Later among the works it cites.
Representations for stable off-policy reinforcement learning
Dibya Ghosh and Marc G Bellemare · 2020
Later among the works it cites.
Rl unplugged: A collection of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Later among the works it cites.
Array programming with numpy
Charles R Harris, K Jarrod Millman, Stéfan J van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2020
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint
Léonard Blier, Corentin Tallec, and Yann Ollivier · 2021
Later among the works it cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, and Sergey Levine · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Clare Lyle, Mark Rowland, Georg Ostrovski, and Will Dabney · 2021
Later among the works it cites.
Dr3: Value-based deep reinforcement learning requires explicit regularization
Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron Courville, George Tucker, and Sergey Levine · 2022
Closest in time.