Fetching the paper…
Reading the bibliography…
Episodic count has been widely used to design a simple yet effective intrinsic motivation for reinforcement learning with a sparse reward.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J. Zico Kolter and Andrew Y. Ng · 2009
Earlier work this paper cites.
Supervised hashing for image retrieval via image representation learning
Rongkai Xia, Yan Pan, Hanjiang Lai, Cong Liu, and Shuicheng Yan · 2014
Earlier work this paper cites.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Earlier work this paper cites.
Count-based exploration in feature space for reinforcement learning
Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt, and Marcus Hutter · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aaron van den Oord, and Remi Munos · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Earlier work this paper cites.
Efficient model–based deep reinforcement learning with variational state tabulation
Dane Corneil, Wulfram Gerstner, and Johanni Brea · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Obstacle tower: A generalization challenge in vision, control, and planning
Arthur Juliani, Ahmed Khalifa, Vincent-Pierre Berges, Jonathan Harper, Hunter Henry, Adam Crespi, Julian Togelius, and Danny Lange · 2019
Cited alongside, same era.
Curiosity-bottleneck: Exploration by distilling task-specific novelty
Youngjin Kim, Wontae Nam, Hyunwoo Kim, Ji-Hoon Kim, and Gunhee Kim · 2019
Cited alongside, same era.
Episodic curiosity through reachability
Learning intrinsic rewards as a bi-level optimization problem
Lunjun Zhang, Bradly C. Stadie, and Jimmy Ba · 2020
Later among the works it cites.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
Dynamic bottleneck for robust self-supervised exploration
Chenjia Bai, Lingxiao Wang, Lei Han, Animesh Garg, Jianye Hao, Peng Liu, and Zhaoran Wang · 2021
Later among the works it cites.
Learning with {amig}o: Adversarially motivated intrinsic goals
Andres Campero, Roberta Raileanu, Heinrich Kuttler, Joshua B. Tenenbaum, Tim Rocktäschel, and Edward Grefenstette · 2021
Later among the works it cites.
First return, then explore
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2021
Later among the works it cites.
Adversarially guided actor-critic
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2019
Cited alongside, same era.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martin Arjovsky, Alexander Pritzel, Andrew Bolt, and Charles Blundell · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2020
Cited alongside, same era.
Seed rl: Scalable and efficient deep-rl with accelerated central inference
Lasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang, and Marcin Michalski · 2020
Cited alongside, same era.
The nethack learning environment
Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Cited alongside, same era.
Count-based exploration with the successor representation
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2020
Cited alongside, same era.
Sample factory: Egocentric 3d control from pixels at 100000 fps with asynchronous reinforcement learning
Aleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav Sukhatme, and Vladlen Koltun · 2020
Cited alongside, same era.
Later among the works it cites.
Neural discrete abstraction of high-dimensional spaces: A case study in reinforcement learning
Petros Giannakopoulos, Aggelos Pikrakis, and Yannis Cotronis · 2021
Later among the works it cites.
Drop-bottleneck: Learning discrete compressed representation for noise-robust exploration
Jaekyeom Kim, Minjung Kim, Dongyeon Woo, and Gunhee Kim · 2021
Later among the works it cites.
Variational intrinsic control revisited
Taehwan Kwon · 2021
Later among the works it cites.
Don’t do what doesn’t matter: Intrinsic motivation with action usefulness
Mathieu Seurin, Florian Strub, Philippe Preux, and Olivier Pietquin · 2021
Later among the works it cites.
Made: Exploration via maximizing deviation from explored regions
Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian, Joseph E. Gonzalez, and Stuart Russell · 2021
Later among the works it cites.
Noveld: A simple yet effective exploration criterion
Tianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu, Kurt Keutzer, Joseph E Gonzalez, and Yuandong Tian · 2021
Later among the works it cites.
Divide and explore: Multi-agent separate exploration with shared intrinsic motivations, 2022
Xiao Jing, Zhenwei Zhu, Hongliang Li, Xin Pei, Yoshua Bengio, Tong Che, and Hongyong Song · 2022
Closest in time.