Fetching the paper…
Reading the bibliography…
Unsupervised pre-training strategies have proven to be highly effective in natural language processing and computer vision.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard E. Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne E. Hubbard, and Lawrence D. Jackel · 1989
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani et al · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael U Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Real analysis for graduate students
Richard F Bass · 2013
Earlier work this paper cites.
Matrix analysis
Rajendra Bhatia · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo J. Rezende · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Earlier work this paper cites.
G. Brockman, Vicki Cheung, Ludwig Pettersson, J. Schneider, John Schulman, Jie Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and P. Abbeel · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and P. Abbeel · 2016
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and P. Abbeel · 2017
Earlier work this paper cites.
Ex2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John D. Co-Reyes, and Sergey Levine · 2017
Earlier work this paper cites.
A laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and P. Abbeel · 2017
Earlier work this paper cites.
Variational option discovery algorithms
Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel · 2018
Earlier work this paper cites.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
John D. Co-Reyes, Yuxuan Liu, Abhishek Gupta, Benjamin Eysenbach, P. Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Eigenoption discovery through the deep successor representation
Marlos C. Machado, Clemens Rosenbaum, Xiaoxiao Guo, Miao Liu, Gerald Tesauro, and Murray Campbell · 2018
Earlier work this paper cites.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun · 2018
Earlier work this paper cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur D. Szlam, and Rob Fergus · 2018
Earlier work this paper cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2019
Earlier work this paper cites.
Self-supervised learning of image embedding for continuous control
Carlos Florensa, Jonas Degrave, Nicolas Manfred Otto Heess, Jost Tobias Springenberg, and Martin A. Riedmiller · 2019
Cited alongside, same era.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric P. Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Wasserstein dependency measure for representation learning
Sherjil Ozair, Corey Lynch, Yoshua Bengio, Aäron van den Oord, Sergey Levine, and Pierre Sermanet · 2019
Cited alongside, same era.
Discovering and achieving goals via world models
Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, and Deepak Pathak · 2021
Later among the works it cites.
Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate
Mirco Mutti, Lorenzo Pratissoli, and Marcello Restelli · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
OpenAI OpenAI, Matthias Plappert, Raul Sampedro, Tao Xu, Ilge Akkaya, Vineet Kosaraju, Peter Welinder, Ruben D’Sa, Arthur Petron, Henrique Pondé de Oliveira Pinto, Alex Paino, Hyeonwoo Noh, Lilian Weng, Qiming Yuan, Casey Chu, and Wojciech Zaremba · 2021
Later among the works it cites.
Information is power: Intrinsic control via information capture
Nick Rhinehart, Jenny Wang, Glen Berseth, John D. Co-Reyes, Danijar Hafner, Chelsea Finn, and Sergey Levine · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, P. Abbeel, and Kimin Lee · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Kumar Gupta · 2019
Cited alongside, same era.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aäron van den Oord, Alexander A. Alemi, and G. Tucker · 2019
Cited alongside, same era.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino J. Gomez · 2019
Cited alongside, same era.
Unsupervised control through non-parametric discriminative rewards
David Warde-Farley, Tom Van de Wiele, Tejas Kulkarni, Catalin Ionescu, Steven Hansen, and Volodymyr Mnih · 2019
Cited alongside, same era.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andew Bolt, and Charles Blundell · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Explore, discover and learn: unsupervised discovery of state-covering skills
Víctor Campos Camúñez, Alex Trott, Caiming Xiong, Richard Socher, Xavier Giró Nieto, and Jordi Torres Viñals · 2020
Cited alongside, same era.
Later among the works it cites.
Learning one representation to optimize all rewards
Ahmed Touati and Yann Ollivier · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Hierarchical reinforcement learning by discovering intrinsic options
Jesse Zhang, Haonan Yu, and Wei Xu · 2021
Later among the works it cites.
Deep hierarchical planning from pixels
Danijar Hafner, Kuang-Huei Lee, Ian S. Fischer, and P. Abbeel · 2022
Later among the works it cites.
Temporal difference learning for model predictive control
Nicklas Hansen, Xiaolong Wang, and Hao Su · 2022
Later among the works it cites.
Wasserstein unsupervised reinforcement learning
Shuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao, and Xiangyang Ji · 2022
Later among the works it cites.
Dropout q-functions for doubly efficient reinforcement learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka · 2022
Later among the works it cites.
Unsupervised skill discovery via recurrent skill training
Zheyuan Jiang, Jingyue Gao, and Jianyu Chen · 2022
Later among the works it cites.
Direct then diffuse: Incremental unsupervised skill discovery for state covering and goal reaching
Pierre-Alexandre Kamienny, Jean Tarbouriech, Alessandro Lazaric, and Ludovic Denoyer · 2022
Later among the works it cites.
Unsupervised reinforcement learning with contrastive intrinsic control
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and P. Abbeel · 2022
Later among the works it cites.
Curiosity-driven exploration via latent bayesian surprise
Pietro Mazzaglia, Ozan Çatal, Tim Verbelen, and B. Dhoedt · 2022
Later among the works it cites.
Lipschitz-constrained unsupervised skill discovery
Seohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee, and Gunhee Kim · 2022
Later among the works it cites.
One after another: Learning incremental skills for a changing world
Nur Muhammad (Mahi) Shafiullah and Lerrel Pinto · 2022
Later among the works it cites.
Learning more skills through optimistic exploration
DJ Strouse, Kate Baumli, David Warde-Farley, Vlad Mnih, and Steven Stenberg Hansen · 2022
Later among the works it cites.
A mixture of surprises for unsupervised reinforcement learning
Andrew Zhao, Matthieu Lin, Yangguang Li, Y. Liu, and Gao Huang · 2022
Later among the works it cites.
Planning goals for exploration
Edward S. Hu, Richard Chang, Oleh Rybkin, and Dinesh Jayaraman · 2023
Closest in time.
Variational curriculum reinforcement learning for unsupervised discovery of skills
Seongun Kim, Kyowoon Lee, and Jaesik Choi · 2023
Closest in time.
Deep laplacian-based options for temporally-extended exploration
Martin Klissarov and Marlos C. Machado · 2023
Closest in time.
Internally rewarded reinforcement learning
Mengdi Li, Xufeng Zhao, Jae Hee Lee, Cornelius Weber, and Stefan Wermter · 2023
Closest in time.
Choreographer: Learning and adapting skills in imagination
Pietro Mazzaglia, Tim Verbelen, B. Dhoedt, Alexandre Lacoste, and Sai Rajeswar · 2023
Closest in time.
Predictable mdp abstraction for unsupervised model-based rl
Seohong Park and Sergey Levine · 2023
Closest in time.
Mastering the unsupervised reinforcement learning benchmark from pixels
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Pich’e, B. Dhoedt, Aaron C. Courville, and Alexandre Lacoste · 2023
Closest in time.
Does zero-shot reinforcement learning exist?
Ahmed Touati, Jérémy Rapin, and Yann Ollivier · 2023
Closest in time.
Optimal goal-reaching reinforcement learning via quasimetric learning
Tongzhou Wang, Antonio Torralba, Phillip Isola, and Amy Zhang · 2023
Closest in time.
Behavior contrastive learning for unsupervised skill discovery
Rushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li, Bin Zhao, Zhen Wang, Peng Liu, and Xuelong Li · 2023
Closest in time.