Fetching the paper…
Reading the bibliography…
In this work we consider partially observable environments with sparse rewards.
State of the art—a survey of partially observable markov decision processes: Theory, models, and algorithms
George E. Monahan · 1982
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Laurenz Wiskott and Terrence J. Sejnowski · 2002
Earlier work this paper cites.
Cholesky factorization of matrices in parallel and ranking of graphs
Dariusz Dereniowski and Kubale Marek · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Satinder Singh, Michael R. James, and Matthew R. Rudary · 2004
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps, 2015
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models, 2015
Bradly C. Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation, 2016
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration, 2016
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
#exploration: A study of count-based exploration for deep reinforcement learning, 2016
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Emi: Exploration with mutual information, 2018
Hyoungseok Kim, Jaekyeom Kim, Yeonwoo Jeong, Sergey Levine, and Hyun Oh Song · 2018
Later among the works it cites.
Episodic curiosity through reachability, 2018
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Unsupervised state representation learning in atari, 2019
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, and R Devon Hjelm · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, Will Dabney, John Quan, and Remi Munos · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Justin Fu, John D. Co-Reyes, and Sergey Levine · 2017
Cited alongside, same era.
Count-based exploration with neural density models, 2017
Georg Ostrovski, Marc G. Bellemare, Aaron van den Oord, and Remi Munos · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Neural predictive belief representations, 2018
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo A. Pires, and Rémi Munos · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Distributed prioritized experience replay, 2018
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Cited alongside, same era.
Large-scale study of curiosity-driven learning, 2018a
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A. Efros
Cited in the paper.
Gradient-based training of slow feature analysis by differentiable approximate whitening, 2019
Merlin Schüler, Hlynur Davíð Hlynsson, and Laurenz Wiskott · 2019
Later among the works it cites.
Whitening and Coloring Batch Transform for GANs
Aliaksandr Siarohin, Enver Sangineto, and Nicu Sebe · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies, 2020
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andew Bolt, and Charles Blundell · 2020
Closest in time.
Whitening for self-supervised representation learning
Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, and Nicu Sebe · 2020
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2020
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Closest in time.
Planning to explore via self-supervised world models, 2020
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Closest in time.