Fetching the paper…
Reading the bibliography…
We introduce the forward-backward (FB) representation of the dynamics of a reward-free Markov decision process.
Finite Markov Chains
J. G. Kemeny and J. L. Snell · 1960
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Introduction to probability
Charles Miller Grinstead and James Laurie Snell · 1997
Earlier work this paper cites.
Markov chains: Gibbs fields, Monte Carlo simulation, and queues
Pierre Brémaud · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Structure in the space of value functions
David Foster and Peter Dayan · 2002
Earlier work this paper cites.
Measure theory
Vladimir I Bogachev · 2007
Earlier work this paper cites.
Proto-value functions: A Laplacian framework for learning representation and control in Markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Dwight Crow, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, David Silver, and Hado P van Hasselt · 2017
Cited alongside, same era.
The hippocampus as a predictive map
Kimberly L Stachenfeld, Matthew M Botvinick, and Samuel J Gershman · 2017
Cited alongside, same era.
Self-correcting models for model-based reinforcement learning
Erik Talvitie · 2017
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, et al · 2018
Later among the works it cites.
Hindsight policy gradients
Paulo Rauber, Avinash Ummadisingu, Filipe Mutz, and Jürgen Schmidhuber · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Unsupervised state representation learning in atari
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, and R Devon Hjelm · 2019
Later among the works it cites.
Disentangled cumulants help successor representations transfer to new tasks
Christopher Grimm, Irina Higgins, Andre Barreto, Denis Teplyashin, Markus Wulfmeier, Tim Hertweck, Raia Hadsell, and Satinder Singh · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning with successor features for navigation across similar environments
Jingwei Zhang, Jost Tobias Springenberg, Joschka Boedecker, and Wolfram Burgard · 2017
Cited alongside, same era.
Universal successor features approximators
Diana Borsa, André Barreto, John Quan, Daniel Mankowitz, Rémi Munos, Hado van Hasselt, David Silver, and Tom Schaul · 2018
Cited alongside, same era.
Modeling the long term future in model-based reinforcement learning
Nan Rosemary Ke, Amanpreet Singh, Ahmed Touati, Anirudh Goyal, Yoshua Bengio, Devi Parikh, and Dhruv Batra · 2018
Cited alongside, same era.
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Steven Hansen, Will Dabney, Andre Barreto, Tom Van de Wiele, David Warde-Farley, and Volodymyr Mnih · 2019
Later among the works it cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint
Léonard Blier, Corentin Tallec, and Yann Ollivier · 2021
Closest in time.
C-learning: Learning to achieve goals via recursive classification
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2021
Closest in time.