Fetching the paper…
Reading the bibliography…
How can a reinforcement learning (RL) agent prepare to solve downstream tasks if those tasks are not known a priori? One approach is unsupervised skill discovery, a class of algorithms that learn a set of policies without access to a reward function.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 1906
Earlier work this paper cites.
The m-center problem
Edward Minieka · 1970
Earlier work this paper cites.
The m-center problem: Minimax facility location
RS Garfinkel, AW Neebe, and MR Rao · 1977
Earlier work this paper cites.
Source coding with side information and universal coding
Robert G Gallager · 1979
Earlier work this paper cites.
Coding of a source with unknown but ordered probabilities
Boris Yakovlevich Ryabko · 1979
Earlier work this paper cites.
Markov decision processes
Martin L Puterman · 1990
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Policy search via density estimation
Andrew Y Ng, Ronald Parr, and Daphne Koller · 1999
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Fundamentals of convex analysis
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal · 2004
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Dual representations for dynamic programming and reinforcement learning
Tao Wang, Michael Bowling, and Dale Schuurmans · 2007
Earlier work this paper cites.
Facility location: concepts, models, algorithms and case studies
Reza Zanjirani Farahani and Masoud Hekmatfar · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Handbook of Markov decision processes: methods and applications , volume 40
Eugene A Feinberg and Adam Shwartz · 2012
Earlier work this paper cites.
Empowerment–an introduction
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
The information geometry of mirror descent
Garvesh Raskutti and Sayan Mukherjee · 2015
Earlier work this paper cites.
Convex analysis
Ralph Tyrell Rockafellar · 2015
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess · 2019
Later among the works it cites.
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham M Kakade, and Sergey Levine · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2019
Later among the works it cites.
Understanding the limitations of variational mutual information estimators
Jiaming Song and Stefano Ermon · 2019
Later among the works it cites.
Flambe: Structural complexity and representation learning of low rank mdps
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Variational option discovery algorithms
Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel · 2018
Cited alongside, same era.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm · 2018
Cited alongside, same era.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
John Co-Reyes, YuXuan Liu, Abhishek Gupta, Benjamin Eysenbach, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil · 2018
Cited alongside, same era.
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Later among the works it cites.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Víctor Campos, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i Nieto, and Jordi Torres · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Steven Hansen, Will Dabney, Andre Barreto, David Warde-Farley, Tom Van de Wiele, and Volodymyr Mnih · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
James Martens · 2020
Later among the works it cites.
Formal limitations on the measurement of mutual information
David McAllester and Karl Stratos · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Reinforcement learning via fenchel-rockafellar duality
Ofir Nachum and Bo Dai · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Yevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao, Dmitry Kalashnikov, Jake Varley, Alex Irpan, Benjamin Eysenbach, Ryan Julian, Chelsea Finn, et al · 2021
Closest in time.
Urlb: Unsupervised reinforcement learning benchmark
Michael Laskin, Denis Yarats, Hao Liu, Kimin Lee, Albert Zhan, Kevin Lu, Catherine Cang, Lerrel Pinto, and Pieter Abbeel · 2021
Closest in time.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Closest in time.
Pretraining representations for data-efficient reinforcement learning
Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, Devon Hjelm, Philip Bachman, and Aaron Courville · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Closest in time.