Fetching the paper…
Reading the bibliography…
We present a new computing model for intrinsic rewards in reinforcement learning that addresses the limitations of existing surprise-driven explorations.
Curious model-building control systems. In Proc. international joint conference on neural networks . 1458–1463
Jürgen Schmidhuber. 1991 · 1991
Earlier work this paper cites.
The nature of emotion: Fundamental questions
Paul Ed Ekman and Richard J Davidson. 1994 · 1994
Earlier work this paper cites.
Bayesian surprise attracts human attention
Laurent Itti and Pierre Baldi. 2005 · 2005
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber. 2010 · 2010
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-Yves Oudeyer. 2012 · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup. 2012 · 2012
Earlier work this paper cites.
Novelty or surprise?
Andrew Barto, Marco Mirolli, and Gianluca Baldassarre. 2013 · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel. 2015 · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos. 2016 · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. 2016 · 2016
Earlier work this paper cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Joshua Achiam and Shankar Sastry. 2017 · 2017
Cited alongside, same era.
Data Efficient Deep Reinforcement Learning through Model-Based Intrinsic Motivation
Mikkel Sannes Nylend. 2017 · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction. In International conference on machine learning . PMLR, 2778–2787
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. 2017 · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
gym-miniworld environment for OpenAI Gym
Maxime Chevalier-Boisvert. 2018 · 2018
Cited alongside, same era.
Agent57: Outperforming the atari human benchmark. In International Conference on Machine Learning . PMLR, 507–517
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell. 2020 · 2020
Later among the works it cites.
SMiRL: Surprise Minimizing Reinforcement Learning in Unstable Environments. In International Conference on Learning Representations
Glen Berseth, Daniel Geng, Coline Manon Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman, and Sergey Levine. 2020 · 2020
Later among the works it cites.
Latent world models for intrinsically motivated exploration
Aleksandr Ermolov and Nicu Sebe. 2020 · 2020
Later among the works it cites.
Overparameterized neural networks implement associative memory
Adityanarayanan Radhakrishnan, Mikhail Belkin, and Caroline Uhler. 2020 · 2020
Later among the works it cites.
Planning to explore via self-supervised world models. In International Conference on Machine Learning . PMLR, 8583–8592
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimalistic Gridworld Environment for OpenAI Gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal. 2018 · 2018
Cited alongside, same era.
Episodic Curiosity through Reachability. In International Conference on Learning Representations
Nikolay Savinov, Anton Raichuk, Damien Vincent, Raphael Marinier, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly. 2018 · 2018
Cited alongside, same era.
Never Give Up: Learning Directed Exploration Strategies. In International Conference on Learning Representations
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martin Arjovsky, Alexander Pritzel, Andrew Bolt, et al · 2019
Cited alongside, same era.
EMI: Exploration with Mutual Information. In International Conference on Machine Learning . PMLR, 3360–3369
Hyoungseok Kim, Jaekyeom Kim, Yeonwoo Jeong, Sergey Levine, and Hyun Oh Song. 2019 · 2019
Cited alongside, same era.
Learning to remember more with less memorization
Hung Le, Truyen Tran, and Svetha Venkatesh. 2019 · 2019
Cited alongside, same era.
Self-supervised exploration via disagreement. In International conference on machine learning . PMLR, 5062–5071
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta. 2019 · 2019
Cited alongside, same era.
Large-Scale Study of Curiosity-Driven Learning. In International Conference on Learning Representations
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros. 2018a
Cited in the paper.
Later among the works it cites.
Model-based episodic memory induces dynamic hybrid controls
Hung Le, Thommen Karimpanal George, Majid Abdolshah, Truyen Tran, and Svetha Venkatesh. 2021 · 2021
Later among the works it cites.
Intrinsic Control of Variational Beliefs in Dynamic Partially-Observed Visual Environments. In ICML 2021 Workshop on Unsupervised Reinforcement Learning
Nicholas Rhinehart, Jenny Wang, Glen Berseth, John D Co-Reyes, Danijar Hafner, Chelsea Finn, and Sergey Levine. 2021 · 2021
Later among the works it cites.
Image Augmentation Based Momentum Memory Intrinsic Reward for Sparse Reward Visual Scenes
Zheng Fang, Biao Zhao, and Guizhong Liu. 2022 · 2022
Later among the works it cites.
Learning to Constrain Policy Optimization with Virtual Trust Region
Hung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen, Kien Do, Sunil Gupta, and Svetha Venkatesh. 2022b · 2022
Later among the works it cites.
Neurocoder: General-purpose computation using stored neural programs. In International Conference on Machine Learning . PMLR, 12204–12221
Hung Le and Svetha Venkatesh. 2022 · 2022
Later among the works it cites.