Fetching the paper…
Reading the bibliography…
What goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the agent should not pursue unobtainable goals.
An algorithm for quadratic programming
Frank, M. and Wolfe, P · 1956
Earlier work this paper cites.
Remarks on some nonparametric estimates of a density function
Rosenblatt, M · 1956
Earlier work this paper cites.
Artificial intelligence and the concept of mind
Newell, A · 1969
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
A frontier-based approach for autonomous exploration
Yamauchi, B · 1997
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
Klyubin, A. S., Polani, D., and Nehaniv, C. L · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M · 2006
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y · 2009
Earlier work this paper cites.
Intrinsically motivated goal exploration for active motor learning in robots: A case study
Baranes, A. and Oudeyer, P.-Y · 2010
Earlier work this paper cites.
Evaluating the efficiency of frontier-based exploration strategies
Holz, D., Basilico, N., Amigoni, F., and Behnke, S · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Elements of information theory
Cover, T. M. and Thomas, J. A · 2012
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, M., Lang, T., Toussaint, M., and Oudeyer, P.-Y · 2012
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Baranes, A. and Oudeyer, P.-Y · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2014
Earlier work this paper cites.
Empowerment–an introduction
Salge, C., Glackin, C., and Polani, D · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fern ′ · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
Thomas, P. S., Theocharous, G., and Ghavamzadeh, M · 2015
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Later among the works it cites.
Meta-learning for semi-supervised few-shot classification
Ren, M., Triantafillou, E., Ravi, S., Snell, J., Swersky, K., Tenenbaum, J. B., Larochelle, H., and Zemel, R. S · 2018
Later among the works it cites.
Model-based active exploration
Shyam, P., Ja ′ · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Density estimation using real nvp
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2016
Cited alongside, same era.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P · 2017
Cited alongside, same era.
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
Curriculum-guided hindsight experience replay
Fang, M., Zhou, T., Du, Y., Han, L., and Zhang, Z · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Later among the works it cites.
Normalizing flows for probabilistic modeling and inference
Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B · 2019
Later among the works it cites.
Protoge: Prototype goal encodings for multi-goal reinforcement learning
Pitis, S., Chan, H., and Ba, J · 2019
Later among the works it cites.
Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2019
Later among the works it cites.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
Trott, A., Zheng, S., Xiong, C., and Socher, R · 2019
Later among the works it cites.
Using a logarithmic mapping to enable lower discount factors in reinforcement learning
Van Seijen, H., Fatemi, M., and Tavakoli, A · 2019
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2019
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
Zhao, R., Sun, X., and Tresp, V · 2019
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., and Blundell, C · 2020
Closest in time.
Leaf: Latent exploration along the frontier
Bharadhwaj, H., Garg, A., and Shkurti, F · 2020
Closest in time.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2020
Closest in time.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
Hartikainen, K., Geng, X., Haarnoja, T., and Levine, S · 2020
Closest in time.
Weakly-supervised reinforcement learning for controllable behavior
Lee, L., Eysenbach, B., Salakhutdinov, R., Finn, C., et al · 2020
Closest in time.
mrl: modular rl
Pitis, S., Chan, H., and Zhao, S · 2020
Closest in time.
Automatic curriculum learning through value disagreement
Zhang, Y., Abbeel, P., and Pinto, L · 2020
Closest in time.