Fetching the paper…
Reading the bibliography…
Acquiring abilities in the absence of a task-oriented reward function is at the frontier of reinforcement learning research.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. J · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Precup, D · 2001
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
Barber, D. and Agakov, F. V · 2003
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Empowerment – an introduction
Salge, C., Glackin, C., and Polani, D · 2014
Earlier work this paper cites.
Robots that can adapt like animals
Cully, A., Clune, J., Tarapore, D., and Mouret, J.-B · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
Mouret, J.-B. and Clune, J · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Learning behavior characterizations for novelty search
Meyerson, E., Lehman, J., and Miikkulainen, R · 2016
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
Pugh, J. K., Soros, L. B., and Stanley, K. O · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Chou, P.-W., Maturana, D., and Scherer, S · 2017
Cited alongside, same era.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S · 2018
Later among the works it cites.
Self-imitation learning
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Later among the works it cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Later among the works it cites.
Universal successor features approximators
Borsa, D., Barreto, A., Quan, J., Mankowitz, D., Munos, R., van Hasselt, H., Silver, D., and Schaul, T · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conti, E., Madhavan, V., Such, F. P., Lehman, J., Stanley, K. O., and Clune, J · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Multi-level discovery of deep options
Fox, R., Krishnan, S., Stoica, I., and Goldberg, K · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Dulac-Arnold, G., et al · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Cited alongside, same era.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Later among the works it cites.
Self-supervised learning of image embedding for continuous control
Florensa, C., Degrave, J., Heess, N., Springenberg, J. T., and Riedmiller, M · 2019
Later among the works it cites.
The minerl competition on sample efficient reinforcement learning using human priors
Guss, W. H., Codel, C., Hofmann, K., Houghton, B., Kuno, N., Milani, S., Mohanty, S., Liebana, D. P., Salakhutdinov, R., Topin, N., et al · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S. M., Singh, K., and Van Soest, A · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2019
Later among the works it cites.
Data-efficient image recognition with contrastive predictive coding
Hénaff, O. J., Razavi, A., Doersch, C., Eslami, S., and Oord, A. v. d · 2019
Later among the works it cites.
Unsupervised curricula for visual meta-reinforcement learning
Jabri, A., Hsu, K., Gupta, A., Eysenbach, B., Levine, S., and Finn, C · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Later among the works it cites.
Competitive experience replay
Liu, H., Trott, A., Socher, R., and Xiong, C · 2019
Later among the works it cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Later among the works it cites.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
Trott, A., Zheng, S., Xiong, C., and Socher, R · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., Van de Wiele, T., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2019
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2020
Closest in time.