Fetching the paper…
Reading the bibliography…
We present BYOL-Explore, a conceptually simple yet general approach for curiosity-driven exploration in visually-complex environments.
Learning to generate artificial fovea trajectories for target detection
Juergen Schmidhuber and Rudolf Huber · 1991
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Instance-based utile distinctions for reinforcement learning with hidden state
R Andrew McCallum · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2004
Earlier work this paper cites.
Intrinsic Motivation for Autonomous Mental Development
Pierre-Yves Oudeyer, Frédéric Kaplan, and Véréna Hafner · 2007
Earlier work this paper cites.
Feature reinforcement learning: Part I. unstructured MDPs
Marcus Hutter et al · 2009
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-Yves Oudeyer · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Q-learning for history-based reinforcement learning
Mayank Daswani, Peter Sunehag, and Marcus Hutter · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Cited alongside, same era.
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, John Quan, Remi Munos, and Will Dabney · 2018
Cited alongside, same era.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
Kristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, and Sergey Levine · 2020
Later among the works it cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Later among the works it cites.
The nethack learning environment
Heinrich Küttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Later among the works it cites.
Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
Silviu Pitis, Harris Chan, Stephen Zhao, Bradly Stadie, and Jimmy Ba · 2020
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual reinforcement learning with imagined goals
Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Cited alongside, same era.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo Avila Pires, Jean-Bastien Grill, Florent Altché, and Rémi Munos · 2019
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Curious: intrinsically motivated modular multi-goal reinforcement learning
Cédric Colas, Pierre Fournier, Mohamed Chetouani, Olivier Sigaud, and Pierre-Yves Oudeyer · 2019
Cited alongside, same era.
Go-explore: a new approach for hard-exploration problems
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2019
Cited alongside, same era.
Shaping belief states with generative environment models for rl
Karol Gregor, Danilo Jimenez Rezende, Frederic Besse, Yan Wu, Hamza Merzic, and Aaron van den Oord · 2019
Cited alongside, same era.
Vitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 2020
Later among the works it cites.
Byol works even without batch statistics
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché, Corentin Tallec, Florian Strub, Andrew Brock, Samuel Smith, Soham De, Razvan Pascanu, Bilal Piot, and Michal Valko · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Active model estimation in markov decision processes
Jean Tarbouriech, Shubhanshu Shekhar, Matteo Pirotta, Mohammad Ghavamzadeh, and Alessandro Lazaric · 2020
Later among the works it cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin Yang, and Russ R Salakhutdinov · 2020
Later among the works it cites.
Provably efficient reward-agnostic navigation with linear value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Automatic curriculum learning through value disagreement
Yunzhi Zhang, Pieter Abbeel, and Lerrel Pinto · 2020
Later among the works it cites.
Near-optimal reward-free exploration for linear mixture mdps with plug-in solver
Xiaoyu Chen, Jiachen Hu, Lin F Yang, and Liwei Wang · 2021
Later among the works it cites.
First return, then explore
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2021
Later among the works it cites.
Geometric entropic exploration
Zhaohan Daniel Guo, Mohammad Gheshlagi Azar, Alaa Saade, Shantanu Thakoor, Bilal Piot, Bernardo Avila Pires, Michal Valko, Thomas Mesnard, Tor Lattimore, and Rémi Munos · 2021
Later among the works it cites.
Podracer architectures for scalable reinforcement learning
Matteo Hessel, Manuel Kroiss, Aidan Clark, Iurii Kemaev, John Quan, Thomas Keck, Fabio Viola, and Hado van Hasselt · 2021
Later among the works it cites.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Broaden your views for self-supervised video learning
Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Pătrăucean, Florent Altché, Michal Valko, et al · 2021
Later among the works it cites.
Understanding self-supervised learning dynamics without contrastive pairs
Yuandong Tian, Xinlei Chen, and Surya Ganguli · 2021
Later among the works it cites.
Reward-free model-based reinforcement learning with linear function approximation
Weitong Zhang, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Near optimal reward-free reinforcement learning
Zihan Zhang, Simon Du, and Xiangyang Ji · 2021
Later among the works it cites.
Large-scale representation learning on graphs via bootstrapping
Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou, Eva L Dyer, Remi Munos, Petar Veličković, and Michal Valko · 2022
Closest in time.
The mechanism of prediction head in non-contrastive self-supervised learning
Zixin Wen and Yuanzhi Li · 2022
Closest in time.