Fetching the paper…
Reading the bibliography…
We study the problem of unsupervised skill discovery, whose goal is to learn a set of diverse and useful skills with no external reward.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C. Pereira, and William Bialek · 2000
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
D. Barber and F. Agakov · 2003
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 2005
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo J. Rezende · 2015
Earlier work this paper cites.
G. Brockman, Vicki Cheung, Ludwig Pettersson, J. Schneider, John Schulman, Jie Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado Van Hasselt, and David Silver · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, T. Lillicrap, and S. Levine · 2017
Earlier work this paper cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell, and Anca Dragan · 2017
Earlier work this paper cites.
A laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel R. D. Rodrigues · 2017
Cited alongside, same era.
Variational option discovery algorithms
Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel · 2018
Cited alongside, same era.
Eigenoption discovery through the deep successor representation
Marlos C. Machado, C. Rosenbaum, Xiaoxiao Guo, Miao Liu, G. Tesauro, and Murray Campbell · 2018
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, R. Józefowicz, Bob McGrew, Jakub W. Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, Jonas Schneider, S. Sidor, Joshua Tobin, P. Welinder, Lilian Weng, and Wojciech Zaremba · 2020
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, P. Sprechmann, Alex Vitvitskyi, Daniel Guo, and C. Blundell · 2020
Later among the works it cites.
Explore, discover and learn: unsupervised discovery of state-covering skills
Víctor Campos Camúñez, Alex Trott, Caiming Xiong, Richard Socher, Xavier Giró Nieto, and Jordi Torres Viñals · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
S. Hansen, Will Dabney, André Barreto, T. Wiele, David Warde-Farley, and V. Mnih · 2020
Later among the works it cites.
Learning to coordinate manipulation skills via skill behavior diversification
Youngwoon Lee, Jingyun Yang, and Joseph J. Lim · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
A PAC-Bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, David A. McAllester, and Nathan Srebro · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba · 2018
Cited alongside, same era.
There is no free lunch in adversarial robustness (but there are unexpected benefits)
D. Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and A. Madry · 2018
Cited alongside, same era.
Challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Daniel Mankowitz, and Todd Hester · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Garage: A toolkit for reproducible reinforcement learning research
The garage contributors · 2019
Cited alongside, same era.
Composing task-agnostic policies with deep reinforcement learning
A. H. Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson, Byron Boots, and Michael C. Yip · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, T. Hubert, K. Simonyan, L. Sifre, Simon Schmitt, A. Guez, Edward Lockhart, Demis Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Variational empowerment as representation learning for goal-conditioned reinforcement learning
Jongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine, and Shixiang Gu · 2021
Later among the works it cites.
Shixiang Shane Gu, Manfred Diaz, Daniel C. Freeman, Hiroki Furuta, Seyed Kamyar Seyed Ghasemipour, Anton Raichuk, Byron David, Erik Frey, Erwin Coumans, and Olivier Bachem · 2021
Later among the works it cites.
Unsupervised skill discovery with bottleneck option learning
Jaekyeom Kim, Seohong Park, and Gunhee Kim · 2021
Later among the works it cites.
Hierarchical reinforcement learning by discovering intrinsic options
Jesse Zhang, Haonan Yu, and Wei Xu · 2021
Later among the works it cites.
Mutual information state intrinsic control
Rui Zhao, Yang Gao, Pieter Abbeel, Volker Tresp, and Wei Xu · 2021
Later among the works it cites.