Fetching the paper…
Reading the bibliography…
Unsupervised reinforcement learning (RL) studies how to leverage environment statistics to learn useful behaviors without the cost of reward engineering.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
What’s interesting?
Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andew Bolt, et al · 2002
Earlier work this paper cites.
All else being equal be empowered
Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv · 2005
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Linda Smith and Michael Gasser · 2005
Earlier work this paper cites.
Metacure: Meta reinforcement learning with empowerment-driven exploration
Jin Zhang, Jianhao Wang, Hao Hu, Tong Chen, Yingfeng Chen, Changjie Fan, and Chongjie Zhang · 2006
Earlier work this paper cites.
The free-energy principle: a rough guide to the brain?
Karl Friston · 2009
Earlier work this paper cites.
Reinforcement learning or active inference?
Karl J Friston, Jean Daunizeau, and Stefan J Kiebel · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Curiosity and boredom based on prediction error as novel internal rewards
Naoyuki Yamamoto and Masumi Ishikawa · 2010
Earlier work this paper cites.
Free-energy minimization and the dark-room problem
Karl Friston, Christopher Thornton, and Andy Clark · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Bebold: Exploration beyond the boundary of explored regions
Tianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu, Kurt Keutzer, Joseph E Gonzalez, and Yuandong Tian · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Empowerment–an introduction
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Maximilian Karl, Justin Bayer, and Patrick van der Smagt · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
Reinforcement learning of pomdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Active inference and learning
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, Giovanni Pezzulo, et al · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Pac reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Joel Z Leibo, Edward Hughes, Marc Lanctot, and Thore Graepel · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2019
Later among the works it cites.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Surprise-based intrinsic motivation for deep reinforcement learning
Joshua Achiam and Shankar Sastry · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Zeming Lin, Ilya Kostrikov, Gabriel Synnaeve, Arthur Szlam, and Rob Fergus · 2017
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Statistical foundation of variational bayes neural networks
Shrijita Bhattacharya and Tapabrata Maiti · 2020
Later among the works it cites.
Learning with amigo: Adversarially motivated intrinsic goals
Andres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B Tenenbaum, Tim Rocktäschel, and Edward Grefenstette · 2020
Later among the works it cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
Michael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre Bayen, Stuart Russell, Andrew Critch, and Sergey Levine · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
Roberta Raileanu and Tim Rocktäschel · 2020
Later among the works it cites.
Continual learning of control primitives: Skill discovery via reset-games
Kelvin Xu, Siddharth Verma, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Efficient online estimation of empowerment for reinforcement learning
Ruihan Zhao, Pieter Abbeel, and Stas Tiomkin · 2020
Later among the works it cites.
Smirl: Surprise minimizing reinforcement learning in unstable environments
Glen Berseth, Daniel Geng, Coline Manon Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman, and Sergey Levine · 2021
Closest in time.
Beyond fine-tuning: Transferring behavior in reinforcement learning
Víctor Campos, Pablo Sprechmann, Steven Stenberg Hansen, Andre Barreto, Steven Kapturowski, Alex Vitvitskyi, Adria Puigdomenech Badia, and Charles Blundell · 2021
Closest in time.
Adversarially guided actor-critic
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist · 2021
Closest in time.
Behavior from the void: Unsupervised active pre-training
Hao Liu and Pieter Abbeel · 2021
Closest in time.
Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate
Mirco Mutti, Lorenzo Pratissoli, and Marcello Restelli · 2021
Closest in time.
Asymmetric self-play for automatic goal discovery in robotic manipulation
OpenAI OpenAI, Matthias Plappert, Raul Sampedro, Tao Xu, Ilge Akkaya, Vineet Kosaraju, Peter Welinder, Ruben D’Sa, Arthur Petron, Henrique Ponde de Oliveira Pinto, et al · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Closest in time.