Fetching the paper…
Reading the bibliography…
In many real-world scenarios, rewards extrinsic to the agent are extremely sparse, or absent altogether.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, Jurgen · 1991
Earlier work this paper cites.
Forward models: Supervised learning with a distal teacher
Jordan, Michael I and Rumelhart, David E · 1992
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Storck, Jan, Hochreiter, Sepp, and Schmidhuber, Jürgen · 1995
Earlier work this paper cites.
An internal model for sensorimotor integration
Wolpert, Daniel M, Ghahramani, Zoubin, and Jordan, Michael I · 1995
Earlier work this paper cites.
Efficient reinforcement learning in factored mdps
Kearns, Michael and Koller, Daphne · 1999
Earlier work this paper cites.
Intrinsic and extrinsic motivations: Classic definitions and new directions
Ryan, Richard; Deci, Edward L · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, Ronen I and Tennenholtz, Moshe · 2002
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
Klyubin, Alexander S, Polani, Daniel, and Nehaniv, Chrystopher L · 2005
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, Satinder P, Barto, Andrew G, and Chentanez, Nuttapong · 2005
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Poupart, Pascal, Vlassis, Nikos, Hoey, Jesse, and Regan, Kevin · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, Pierre-Yves, Kaplan, Frdric, and Hafner, Verena V · 2007
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, Pierre-Yves and Kaplan, Frederic · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, Jürgen · 2010
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Sun, Yi, Gomez, Faustino, and Schmidhuber, Jürgen · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, Manuel, Lang, Tobias, Toussaint, Marc, and Oudeyer, Pierre-Yves · 2012
Earlier work this paper cites.
Curiosity and motivation
Silvia, Paul J · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Still, Susanne and Precup, Doina · 2012
Cited alongside, same era.
Learning and exploration in action-perception loops
Little, Daniel Y and Sommer, Friedrich T · 2014
Cited alongside, same era.
Learning to see by moving
Agrawal, Pulkit, Carreira, Joao, and Malik, Jitendra · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, Djork-Arné, Unterthiner, Thomas, and Hochreiter, Sepp · 2015
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
Doersch, Carl, Gupta, Abhinav, and Efros, Alexei A · 2015
Cited alongside, same era.
Learning to act by predicting the future
Dosovitskiy, Alexey and Koltun, Vladlen · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Houthooft, Rein, Chen, Xi, Duan, Yan, Schulman, John, De Turck, Filip, and Abbeel, Pieter · 2016
Later among the works it cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Kempka, Michał, Wydmuch, Marek, Runc, Grzegorz, Toczek, Jakub, and Jaśkowski, Wojciech · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goroshin, Ross, Bruna, Joan, Tompson, Jonathan, Eigen, David, and LeCun, Yann · 2015
Cited alongside, same era.
Learning image representations tied to ego-motion
Jayaraman, Dinesh and Grauman, Kristen · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, Shakir and Rezende, Danilo Jimenez · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, Junhyuk, Guo, Xiaoxiao, Lee, Honglak, Lewis, Richard L, and Singh, Satinder · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, Bradly C, Levine, Sergey, and Abbeel, Pieter · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
Wang, Xiaolong and Gupta, Abhinav · 2015
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Van Roy, Benjamin · 2016
Later among the works it cites.
Super mario bros. in openai gym
Paquette, Philip · 2016
Later among the works it cites.
Context encoders: Feature learning by inpainting
Pathak, Deepak, Krahenbuhl, Philipp, Donahue, Jeff, Darrell, Trevor, and Efros, Alexei A · 2016
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, Haoran, Houthooft, Rein, Foote, Davis, Stooke, Adam, Chen, Xi, Duan, Yan, Schulman, John, De Turck, Filip, and Abbeel, Pieter · 2016
Later among the works it cites.
Ex2: Exploration with exemplar models for deep reinforcement learning
Fu, Justin, Co-Reyes, John D, and Levine, Sergey · 2017
Closest in time.
Variational intrinsic control
Gregor, Karol, Rezende, Danilo Jimenez, and Wierstra, Daan · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2017
Closest in time.
Learning to navigate in complex environments
Mirowski, Piotr, Pascanu, Razvan, Viola, Fabio, Soyer, Hubert, Ballard, Andy, Banino, Andrea, Denil, Misha, Goroshin, Ross, Sifre, Laurent, Kavukcuoglu, Koray, et al · 2017
Closest in time.
Loss is its own reward: Self-supervision for reinforcement learning
Shelhamer, Evan, Mahmoudieh, Parsa, Argus, Max, and Darrell, Trevor · 2017
Closest in time.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, Sainbayar, Kostrikov, Ilya, Szlam, Arthur, and Fergus, Rob · 2017
Closest in time.