Fetching the paper…
Reading the bibliography…
One of the key capabilities of intelligent agents is the ability to discover useful skills without external supervision.
The IM algorithm: a variational approach to information maximization
Barber, D. and Agakov, F · 2003
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
Estimating mutual information
Kraskov, A., Stögbauer, H., and Grassberger, P · 2004
Earlier work this paper cites.
Modern control engineering
Ogata, K. et al · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mutual information state intrinsic control
Zhao, R., Gao, Y., Abbeel, P., Tresp, V., and Xu, W · 2017
Earlier work this paper cites.
Variational option discovery algorithms
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P · 2018
Earlier work this paper cites.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
Co-Reyes, J. D., Liu, Y., Gupta, A., Eysenbach, B., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., Kumar, V., and Zaremba, W · 2018
Earlier work this paper cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Kostrikov, I., Szlam, A. D., and Fergus, R · 2018
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A. J., and Klimov, O · 2019
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Cited alongside, same era.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Cited alongside, same era.
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A. K · 2019
Cited alongside, same era.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2019
Cited alongside, same era.
Discovering and achieving goals via world models
Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
OpenAI, O., Plappert, M., Sampedro, R., Xu, T., Akkaya, I., Kosaraju, V., Welinder, P., D’Sa, R., Petron, A., de Oliveira Pinto, H. P., Paino, A., Noh, H., Weng, L., Yuan, Q., Chu, C., and Zaremba, W · 2021
Later among the works it cites.
Learning one representation to optimize all rewards
Touati, A. and Ollivier, Y · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Skill-based reinforcement learning with intrinsic reward matching
Adeniji, A., Xie, A., and Abbeel, P · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Campos Camúñez, V., Trott, A., Xiong, C., Socher, R., Giró Nieto, X., and Torres Viñals, J · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T. P., Ba, J., and Norouzi, M · 2020
Cited alongside, same era.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Wiele, T., Warde-Farley, D., and Mnih, V · 2020
Cited alongside, same era.
Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
Pitis, S., Chan, H., Zhao, S., Stadie, B. C., and Ba, J · 2020
Cited alongside, same era.
Skew-Fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2020
Cited alongside, same era.
Composing task-agnostic policies with deep reinforcement learning
Qureshi, A. H., Johnson, J. J., Qin, Y., Henderson, T., Boots, B., and Yip, M. C · 2020
Cited alongside, same era.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D · 2020
Cited alongside, same era.
It takes four to tango: Multiagent selfplay for automatic curriculum generation
Du, Y., Abbeel, P., and Grover, A · 2022
Later among the works it cites.
Unsupervised skill discovery via recurrent skill training
Jiang, Z., Gao, J., and Chen, J · 2022
Later among the works it cites.
Direct then diffuse: Incremental unsupervised skill discovery for state covering and goal reaching
Kamienny, P.-A., Tarbouriech, J., Lazaric, A., and Denoyer, L · 2022
Later among the works it cites.
Unsupervised reinforcement learning with contrastive intrinsic control
Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P · 2022
Later among the works it cites.
Lipschitz-constrained unsupervised skill discovery
Park, S., Choi, J., Kim, J., Lee, H., and Kim, G · 2022
Later among the works it cites.
Unsupervised model-based pre-training for data-efficient control from pixels
Rajeswar, S., Mazzaglia, P., Verbelen, T., Pich’e, A., Dhoedt, B., Courville, A. C., and Lacoste, A · 2022
Later among the works it cites.
Masked world models for visual control
Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., and Abbeel, P · 2022
Later among the works it cites.
One after another: Learning incremental skills for a changing world
Shafiullah, N. M. M. and Pinto, L · 2022
Later among the works it cites.
Learning more skills through optimistic exploration
Strouse, D., Baumli, K., Warde-Farley, D., Mnih, V., and Hansen, S. S · 2022
Later among the works it cites.
Does zero-shot reinforcement learning exist?
Touati, A., Rapin, J., and Ollivier, Y · 2022
Later among the works it cites.
A mixture of surprises for unsupervised reinforcement learning
Zhao, A., Lin, M., Li, Y., Liu, Y., and Huang, G · 2022
Later among the works it cites.