Fetching the paper…
Reading the bibliography…
Learning to reach goal states and learning diverse skills through mutual information (MI) maximization have been proposed as principled frameworks for self-supervised reinforcement learning, allowing agents to acquire broadly applicable multitask policies with minimal reward engineering.
Curious model-building control systems
Schmidhuber, J · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
Barber, D. and Agakov, F. V · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
All else being equal be empowered
Klyubin, A. S., Polani, D., and Nehaniv, C. L · 2005
Earlier work this paper cites.
A new view of automatic relevance determination
Wipf, D. P. and Nagarajan, S. S · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Empowerment for continuous agent—environment systems
Jung, T., Polani, D., and Stone, P · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Empowerment – An Introduction
Salge, C., Glackin, C., and Polani, D · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Earlier work this paper cites.
Universal Value Function Approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Hindsight Experience Replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
TF-Agents: A library for reinforcement learning in tensorflow
Guadarrama, S., Korattikara, A., Ramirez, O., Castro, P., Holly, E., Fishman, S., Wang, K., Gonina, E., Wu, N., Kokiopoulou, E., Sbaiz, L., Smith, J., Bartók, G., Berent, J., Harris, C., Vanhoucke, V., and Brevdo, E · 2018
Later among the works it cites.
Unsupervised meta-learning for reinforcement learning
Gupta, A., Eysenbach, B., Finn, C., and Levine, S · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Successor Features for Transfer in Reinforcement Learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Silver, D., and van Hasselt, H · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2017
Cited alongside, same era.
Variational Intrinsic Control
Gregor, K., Rezende, D. J., and Wierstra, D · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Cited alongside, same era.
Hazan, E., Kakade, S. M., Singh, K., and Van Soest, A · 2018
Later among the works it cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Later among the works it cites.
The laplacian in rl: Learning representations with efficient approximations
Wu, Y., Tucker, G., and Nachum, O · 2018
Later among the works it cites.
Diversity is All You Need: Learning Skills without a Reward Function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Ghasemipour, S. K. S., Zemel, R., and Gu, S · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Later among the works it cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S · 2019
Later among the works it cites.
Unsupervised Control Through Non-Parametric Discriminative Rewards
Warde-Farley, D., Van de Wiele, T., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2019
Later among the works it cites.
Fast Task Inference with Variational Intrinsic Successor Features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2020
Later among the works it cites.