Fetching the paper…
Reading the bibliography…
We introduce Contrastive Intrinsic Control (CIC), an algorithm for unsupervised skill discovery that maximizes the mutual information between state-transitions and latent skill vectors.
Nonparametric entropy estimation: An overview
Beirlant, J · 1997
Earlier work this paper cites.
The im algorithm: A variational approach to information maximization
Barber, D. and Agakov, F. V · 2003
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Singh, H., Misra, N., Hnizdo, V., Fedorowicz, A., and Demchuk, E · 2003
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Van Hasselt, H., and Silver, D · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Earlier work this paper cites.
Variational intrinsic control
Gregor, K., Rezende, D. J., and Wierstra, D · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Variational option discovery algorithms
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P · 2018
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Learning dexterous in-hand manipulation
OpenAI · 2018
Cited alongside, same era.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M · 2018
Cited alongside, same era.
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Later among the works it cites.
On variational bounds of mutual information
Poole, B., Ozair, S., van den Oord, A., Alemi, A., and Tucker, G · 2019
Later among the works it cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Warde-Farley, D., de Wiele, T. V., and Mnih, V · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Unsupervised control through non-parametric discriminative rewards, 2018
Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2018
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Cited alongside, same era.
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2020
Later among the works it cites.
On mutual information maximization for representation learning
Tschannen, M., Djolonga, J., Rubenstein, P. K., Gelly, S., and Lucic, M · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice, 2021
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
Variational intrinsic control revisited
Kwon, T · 2021
Later among the works it cites.
Urlb: Unsupervised reinforcement learning benchmark, 2021
Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P · 2021
Later among the works it cites.
Unsupervised learning for reinforcement learning, 2021
Srinivas, A. and Abbeel, P · 2021
Later among the works it cites.
Learning more skills through optimistic exploration
Strouse, D., Baumli, K., Warde-Farley, D., Mnih, V., and Hansen, S · 2021
Later among the works it cites.
Discovering a set of policies for the worst case reward, 2021
Zahavy, T., Barreto, A., Mankowitz, D. J., Hou, S., O’Donoghue, B., Kemaev, I., and Singh, S. B · 2021
Later among the works it cites.