Fetching the paper…
Reading the bibliography…
Having the ability to acquire inherent skills from environments without any external rewards or supervision like humans is an important problem.
Emergence of invariance and disentanglement in deep representations
Achille, A. and Soatto, S · 1980
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 2000
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I., and Frey, B · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., and Dragan, A · 2017
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Variational option discovery algorithms
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P · 2018
Cited alongside, same era.
Discovering interpretable representations for both deep generative and discriminative models
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D., and Hester, T · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Later among the works it cites.
Garage: A toolkit for reproducible reinforcement learning research
garage contributors, T · 2019
Later among the works it cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Schölkopf, B., and Bachem, O · 2019
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adel, T., Ghahramani, Z., and Weller, A · 2018
Cited alongside, same era.
Isolating sources of disentanglement in variational autoencoders
Chen, R. T., Li, X., Grosse, R. B., and Duvenaud, D. K · 2018
Cited alongside, same era.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
Co-Reyes, J. D., Liu, Y., Gupta, A., Eysenbach, B., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation
Kahn, G., Villaflor, A., Ding, B., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Disentangling by factorising
Kim, H. and Mnih, A · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2018
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2018
Cited alongside, same era.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Model-based active exploration
Shyam, P., Jaśkowski, W., and Gomez, F · 2019
Later among the works it cites.
Robel: Robotics benchmarks for learning with low-cost robots
Ahn, M., Zhu, H., Hartikainen, K., Ponte, H., Gupta, A., Levine, S., and Kumar, V · 2020
Later among the works it cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2020
Later among the works it cites.
Explore, discover and learn: unsupervised discovery of state-covering skills
Campos Camúñez, V., Trott, A., Xiong, C., Socher, R., Giró Nieto, X., and Torres Viñals, J · 2020
Later among the works it cites.
Theory and evaluation metrics for learning disentangled representations
Do, K. and Tran, T · 2020
Later among the works it cites.
Hierarchical reinforcement learning by discovering intrinsic options
Zhang, J., Yu, H., and Xu, W · 2021
Closest in time.