Fetching the paper…
Reading the bibliography…
In Multi-Goal Reinforcement Learning, an agent learns to achieve multiple goals with a goal-conditioned policy.
Weighted entropy
Guiaşu, S · 1971
Earlier work this paper cites.
Inequalities: theory of majorization and its applications , volume 143
Marshall, A. W., Olkin, I., and Arnold, B. C · 1979
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Hierarchical learning in stochastic domains: Preliminary results
Kaelbling, L. P · 1993
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning based on subgoal discovery and subpolicy specialization
Bakker, B. and Schmidhuber, J · 2004
Earlier work this paper cites.
Inequalities of karamata, schur and muirhead, and some applications
Kadelburg, Z., Dukic, D., Lukic, M., and Matic, I · 2005
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Ng, A. Y., Coates, A., Diel, M., Ganapathi, V., Schulte, J., Tse, B., Berger, E., and Liang, E · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
Pearson correlation coefficient
Benesty, J., Chen, J., Huang, Y., and Cohen, I · 2009
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J., Yang, Q., et al · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Machine learning: A probabilistic perspective. adaptive computation and machine learning, 2012
Murphy, K. P · 2012
Earlier work this paper cites.
Universal option models
Szepesvari, C., Sutton, R. S., Modayil, J., Bhatnagar, S., et al · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Tensorflow: a system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
An alternative softmax operator for reinforcement learning
Asadi, K. and Littman, M. L · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Weighted entropy: basic inequalities
Kelbert, M., Stuhl, I., and Suhov, Y · 2017
Later among the works it cites.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Later among the works it cites.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., and Hinton, G · 2017
Later among the works it cites.
Learning to push by grasping: Using multiple tasks for effective learning
Pinto, L. and Gupta, A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bordes, A., Boureau, Y.-L., and Weston, J · 2016
Cited alongside, same era.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Improving policy gradient by exploring under-appreciated rewards
Nachum, O., Norouzi, M., and Schuurmans, D · 2016
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Cited alongside, same era.
Rauber, P., Mutz, F., and Schmidhuber, J · 2017
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P · 2017
Later among the works it cites.
Two-stream rnn/cnn for action recognition in 3d videos
Zhao, R., Ali, H., and Van der Smagt, P · 2017
Later among the works it cites.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Later among the works it cites.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
Péré, A., Forestier, S., Sigaud, O., and Oudeyer, P.-Y · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Later among the works it cites.
Efficient dialog policy learning via positive memory retention
Zhao, R. and Tresp, V · 2018
Later among the works it cites.
Learning goal-oriented visual dialog via tempered policy gradient
Zhao, R. and Tresp, V · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Closest in time.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., Van de Wiele, T., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2019
Closest in time.
Curiosity-driven experience prioritization via density estimation
Zhao, R. and Tresp, V · 2019
Closest in time.