Fetching the paper…
Reading the bibliography…
In reinforcement learning, temporal difference-based algorithms can be sample-inefficient: for instance, with sparse rewards, no learning occurs until a reward is observed.
On the asymptotic distribution of the eigenvalues and eigenfunctions of elliptic differential operators
Lars Gårding · 1953
Earlier work this paper cites.
Finite Markov Chains
J. G. Kemeny and J. L. Snell · 1960
Earlier work this paper cites.
An improved newton iteration for the generalized inverse of a matrix, with applications
Victor Pan and Robert Schreiber · 1991
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Logarithmic Sobolev inequalities for finite Markov chains
Persi Diaconis and Laurent Saloff-Coste · 1996
Earlier work this paper cites.
Introduction to probability
Charles Miller Grinstead and James Laurie Snell · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Markov chains: Gibbs fields, Monte Carlo simulation, and queues , volume 31
Pierre Brémaud · 1999
Earlier work this paper cites.
A panoramic view of Riemannian geometry
Marcel Berger · 2003
Earlier work this paper cites.
Inequalities for the L1 deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Probability measures on metric spaces , volume 352
Kalyanapuram R Parthasarathy · 2005
Earlier work this paper cites.
Ergodic properties of Markov processes
Martin Hairer · 2006
Earlier work this paper cites.
Convergence of Markov processes
Martin Hairer · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
Dynamic Programming and Optimal Control , volume 2
Dimitri P. Bertsekas · 2012
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Cited alongside, same era.
Hindsight experience replay
Universal successor representations for transfer reinforcement learning
Chen Ma, Junfeng Wen, and Yoshua Bengio · 2018
Later among the works it cites.
Eigenoption discovery through the deep successor representation
Marlos C. Machado, Clemens Rosenbaum, Xiaoxiao Guo, Miao Liu, Gerald Tesauro, and Murray Campbell · 2018
Later among the works it cites.
Approximate temporal difference learning is a gradient descent for reversible policies, 2018
Yann Ollivier · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Geometric insights into the convergence of nonlinear td learning
David Brandfonbrener and Joan Bruna · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Andre Barreto, Will Dabney, Remi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Cited alongside, same era.
Advantages and limitations of using successor features for transfer in reinforcement learning
Lucas Lehnert, Stefanie Tellex, and Michael L Littman · 2017
Cited alongside, same era.
Learning to push by grasping: Using multiple tasks for effective learning
Lerrel Pinto and Abhinav Gupta · 2017
Cited alongside, same era.
The hippocampus as a predictive map
Kimberly L Stachenfeld, Matthew M Botvinick, and Samuel J Gershman · 2017
Cited alongside, same era.
On the weyl’s law for discretized elliptic operators
Jinchao Xu, Hongxuan Zhang, and Ludmil Zikatanov · 2017
Cited alongside, same era.
Deep reinforcement learning with successor features for navigation across similar environments
J. Zhang, J. T. Springenberg, J. Boedecker, and W. Burgard · 2017
Cited alongside, same era.
The paths perspective on value learning
Sam Greydanus and Chris Olah · 2019
Later among the works it cites.
Count-based exploration with the successor representation, 2019
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2019
Later among the works it cites.
Value function estimation in Markov reward processes: Instance-dependent ℓ ∞ \ell_{\infty} -bounds for policy evaluation
Ashwin Pananjady and Martin J. Wainwright · 2019
Later among the works it cites.
Model-based RL in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
André Barreto, Shaobo Hou, Diana Borsa, David Silver, and Doina Precup · 2020
Later among the works it cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Later among the works it cites.
Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel, David Silver, André Barreto, and Diana Borsa · 2020
Later among the works it cites.
Weyl law
Wikipedia · 2021
Closest in time.