Fetching the paper…
Reading the bibliography…
Goal-conditioned policies are used in order to break down complex reinforcement learning (RL) problems by using subgoals, which can be defined either in state space or in a latent feature space.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Modern multidimensional scaling: Theory and applications
Borg, I. and Groenen, P · 2003
Earlier work this paper cites.
On spectral graph drawing
Koren, Y · 2003
Earlier work this paper cites.
A novel way of computing similarities between nodes of a graph, with application to collaborative recommendation
Fouss, F., Pirotte, A., and Saerens, M · 2005
Earlier work this paper cites.
Random-walk computation of similarities between nodes of a graph with application to collaborative recommendation
Fouss, F., Pirotte, A., Renders, J.-M., and Saerens, M · 2007
Earlier work this paper cites.
A tutorial on spectral clustering
Luxburg, U · 2007
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P · 2017
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Nair, A., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Later among the works it cites.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
Péré, A., Forestier, S., Sigaud, O., and Oudeyer, P.-Y · 2018
Later among the works it cites.
Semi-parametric topological memory for navigation
Savinov, N., Dosovitskiy, A., and Koltun, V · 2018
Later among the works it cites.
Learning goal embeddings via self-play for hierarchical reinforcement learning
Sukhbaatar, S., Denton, E., Szlam, A., and Fergus, R · 2018
Later among the works it cites.
Learning actionable representations with goal conditioned policies
Ghosh, D., Gupta, A., and Levine, S · 2019
Closest in time.
Hindsight policy gradients
Rauber, P., Ummadisingu, A., Mutz, F., and Schmidhuber, J · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Autonomous task sequencing for customized curriculum design in reinforcement learning
Narvekar, S., Sinapov, J., and Stone, P · 2017
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Cited alongside, same era.
Closest in time.
Episodic curiosity through reachability
Savinov, N., Raichuk, A., Vincent, D., Marinier, R., Pollefeys, M., Lillicrap, T., and Gelly, S · 2019
Closest in time.
The laplacian in RL: Learning representations with efficient approximations
Wu, Y., Tucker, G., and Nachum, O · 2019
Closest in time.