Fetching the paper…
Reading the bibliography…
Unsupervised skill learning aims to learn a rich repertoire of behaviors without external supervision, providing artificial agents with the ability to control and influence the environment.
Learning latent plans from play, 2019
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 1903
Earlier work this paper cites.
Efficient exploration via state marginal matching, 2019
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 1906
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2, 2019
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 1906
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning, 2019
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 1910
Earlier work this paper cites.
Solving rubik’s cube with a robot hand, 2019
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 1910
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 1910
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Object recognition with gradient-based learning
Yann LeCun, Patrick Haffner, Léon Bottou, and Yoshua Bengio · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Doina Precup · 2000
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Harshinder Singh, Neeraj Misra, Vladimir Hnizdo, Adam Fedorowicz, and Eugene Demchuk · 2003
Earlier work this paper cites.
k-means++: The advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2006
Earlier work this paper cites.
Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate, 2020
Mirco Mutti, Lorenzo Pratissoli, and Marcello Restelli · 2007
Earlier work this paper cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning, 2020
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum · 2010
Earlier work this paper cites.
Accelerating reinforcement learning with learned skill priors, 2020
Karl Pertsch, Youngwoon Lee, and Joseph J. Lim · 2010
Earlier work this paper cites.
Parrot: Data-driven behavioral priors for reinforcement learning, 2020
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2011
Earlier work this paper cites.
Latent skill planning for exploration and transfer, 2020
Kevin Xie, Homanga Bharadhwaj, Danijar Hafner, Animesh Garg, and Florian Shkurti · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation, 2013
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Auto-encoding variational bayes, 2013
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Changing the environment based on empowerment as intrinsic motivation
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Variational intrinsic control, 2016
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, and Marc Lanctot · 2016
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Discovering motor programs by recomposing demonstrations
Tanmay Shankar, Shubham Tulsiani, Lerrel Pinto, and Abhinav Gupta · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
Perception, action, and intrinsic motivation in infants’ motor-skill development
Daniela Corbetta · 2021
Later among the works it cites.
The information geometry of unsupervised reinforcement learning, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu · 2017
Cited alongside, same era.
Variational option discovery algorithms
Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, P. Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Fast decoding in sequence models using discrete latent variables, 2018
Łukasz Kaiser, Aurko Roy, Ashish Vaswani, Niki Parmar, Samy Bengio, Jakob Uszkoreit, and Noam Shazeer · 2018
Cited alongside, same era.
Compile: Compositional imitation learning and execution, 2018
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2018
Cited alongside, same era.
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2021
Later among the works it cites.
Hierarchical skills for efficient exploration, 2021
Jonas Gehring, Gabriel Synnaeve, Andreas Krause, and Nicolas Usunier · 2021
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
Entropic desired dynamics for intrinsic control
Steven Stenberg Hansen, Guillaume Desjardins, Kate Baumli, David Warde-Farley, Nicolas Heess, Simon Osindero, and Volodymyr Mnih · 2021
Later among the works it cites.
Unsupervised skill discovery with bottleneck option learning
Jaekyeom Kim, Seohong Park, and Gunhee Kim · 2021
Later among the works it cites.
Urlb: Unsupervised reinforcement learning benchmark, 2021
Michael Laskin, Denis Yarats, Hao Liu, Kimin Lee, Albert Zhan, Kevin Lu, Catherine Cang, Lerrel Pinto, and Pieter Abbeel · 2021
Later among the works it cites.
Curiosity-driven exploration via latent bayesian surprise
Pietro Mazzaglia, Ozan Çatal, Tim Verbelen, and B. Dhoedt · 2021
Later among the works it cites.
Unsupervised skill-discovery and skill-learning in minecraft, 2021
Juan José Nieto, Roger Creus, and Xavier Giro-i Nieto · 2021
Later among the works it cites.
Interesting object, curious agent: Learning task-agnostic exploration, 2021
Simone Parisi, Victoria Dean, Deepak Pathak, and Abhinav Gupta · 2021
Later among the works it cites.
Demonstration-guided reinforcement learning with learned skills, 2021
Karl Pertsch, Youngwoon Lee, Yue Wu, and Joseph J. Lim · 2021
Later among the works it cites.
Unsupervised learning for reinforcement learning
Aravind Srinivas and Pieter Abbeel · 2021
Later among the works it cites.
Skill preferences: Learning to extract and execute robotic skills from human feedback, 2021
Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
General policy evaluation and improvement by learning to identify few but crucial states, 2022
Francesco Faccio, Aditya Ramesh, Vincent Herrmann, Jean Harb, and Jürgen Schmidhuber · 2022
Closest in time.
Deep hierarchical planning from pixels, 2022
Danijar Hafner, Kuang-Huei Lee, Ian Fischer, and Pieter Abbeel · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery, 2022
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
Lipschitz-constrained unsupervised skill discovery, 2022
Seohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee, and Gunhee Kim · 2022
Closest in time.
Unsupervised model-based pre-training for data-efficient control from pixels, 2022
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché, Bart Dhoedt, Aaron Courville, and Alexandre Lacoste · 2022
Closest in time.
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning, 2022
Laura Smith, Ilya Kostrikov, and Sergey Levine · 2022
Closest in time.