Fetching the paper…
Reading the bibliography…
Goal-conditioned reinforcement learning (GCRL) has a wide range of potential real-world applications, including manipulation and navigation problems in robotics.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. 1993 · 1993
Earlier work this paper cites.
Predictive representations of state
Littman, M.; and Sutton, R. S. 2001 · 2001
Earlier work this paper cites.
An inductive bias for distances: Neural nets that respect the triangle inequality
Pitis, S.; Chan, H.; Jamali, K.; and Ba, J. 2020 · 2002
Earlier work this paper cites.
Learning predictive state representations
Singh, S. P.; Littman, M. L.; Jong, N. K.; Pardoe, D.; and Stone, P. 2003 · 2003
Earlier work this paper cites.
Value function approximation with diffusion wavelets and laplacian eigenfunctions
Mahadevan, S.; and Maggioni, M. 2005 · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T.; Horgan, D.; Gregor, K.; and Silver, D. 2015 · 2015
Cited alongside, same era.
Deep successor reinforcement learning
Kulkarni, T. D.; Saeedi, A.; Gautam, S.; and Gershman, S. J. 2016 · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Pieter Abbeel, O.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
Universal successor features approximators
Borsa, D.; Barreto, A.; Quan, J.; Mankowitz, D.; Munos, R.; Van Hasselt, H.; Silver, D.; and Schaul, T. 2018 · 2018
Soft actor-critic algorithms and applications
Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. 2018 · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M.; Andrychowicz, M.; Ray, A.; McGrew, B.; Baker, B.; Powell, G.; Schneider, J.; Tobin, J.; Chociej, M.; Welinder, P.; et al. 2018 · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
Hong, Z.-W.; Yang, G.; and Agrawal, P. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S.; Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Wang, T.; and Isola, P. 2022 · 2022
Closest in time.