Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), agents often operate in partially observed and uncertain environments.
On a measure of the information provided by an experiment
D. V. Lindley · 1956
Earlier work this paper cites.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments, 1990
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jean-Arcady Meyer and Stewart W. Wilson · 1991
Earlier work this paper cites.
Efficient exploration in reinforcement learning, 1992
Sebastian B. Thrun · 1992
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
Christopher G. Atkeson and Juan Carlos Santamaria · 1997
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
J. Beirlant, E. J. Dudewicz, and L. Györfi · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G. Barto, and Satinder P. Singh · 2005
Earlier work this paper cites.
Model Predictive Control
Eduardo F. Camacho and Carlos Bordons Alba · 2007
Earlier work this paper cites.
Predictive coding under the free-energy principle
Karl Friston and Stefan Kiebel · 2009
Earlier work this paper cites.
Bayesian surprise attracts human attention
Laurent Itti and Pierre Baldi · 2009
Earlier work this paper cites.
The free-energy principle: a unified brain theory?
Karl Friston · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. P. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Juergen Schmidhuber · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-yves Oudeyer · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Whatever next? predictive brains, situated agents, and the future of cognitive science
Andy Clark · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Active inference and epistemic value
Karl Friston, Francesco Rigoli, Dimitri Ognibene, Christoph Mathys, Thomas Fitzgerald, and Giovanni Pezzulo · 2015
Earlier work this paper cites.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Gaussian processes for data-efficient learning in robotics and control
Marc Peter Deisenroth, Dieter Fox, and Carl Edward Rasmussen · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Earlier work this paper cites.
Uncertainty and exploration in a restless bandit problem
Maarten Speekenbrink and Emmanouil Konstantinidis · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C. Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Cited alongside, same era.
VIME: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
Deep active inference
Kai Ueltzhöffer · 2018
Later among the works it cites.
The discrete and continuous brain: From decisions to movement—and back again
Thomas Parr and Karl J. Friston · 2018
Later among the works it cites.
Information maximizing exploration with a latent dynamics model
Trevor Barron, Oliver Obst, and Heni Ben Amor · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Tim Pearce, Nicolas Anastassacos, Mohamed Zaki, and Andy Neely · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maximilian Karl, Maximilian Soelch, Justin Bayer, and Patrick van der Smagt · 2016
Cited alongside, same era.
Improving PILCO with bayesian neural network dynamics models
Yarin Gal, Rowan McAllister, and Carl Edward Rasmussen · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Active inference: A process theory
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, and Giovanni Pezzulo · 2017
Cited alongside, same era.
Active inference, curiosity and insight
Karl J. Friston, Marco Lin, Christopher D. Frith, Giovanni Pezzulo, J. Allan Hobson, and Sasha Ondobaka · 2017
Cited alongside, same era.
Active inference, curiosity and insight
Karl J. Friston, Marco Lin, Christopher D. Frith, Giovanni Pezzulo, J. Allan Hobson, and Sasha Ondobaka · 2017
Cited alongside, same era.
Variational inference: A review for statisticians
David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe · 2017
Cited alongside, same era.
Approximate bayesian inference in spatial environments
Atanas Mirchev, Baris Kayalibay, Maximilian Soelch, Patrick van der Smagt, and Justin Bayer · 2018
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 2019
Closest in time.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Closest in time.
A free energy principle for a particular physics
Karl Friston · 2019
Closest in time.
Deep active inference as variational policy gradients
Beren Millidge · 2019
Closest in time.
Bayesian policy selection using active inference
Ozan Çatal, Johannes Nauta, Tim Verbelen, Pieter Simoens, and Bart Dhoedt · 2019
Closest in time.
Pid control as a process of active inference with linear generative models
Manuel Baltieri and Christopher L Buckley · 2019
Closest in time.
Bootstrapping the expressivity with model-based planning
Kefan Dong, Yuping Luo, and Tengyu Ma · 2019
Closest in time.
Learning action-oriented models through active inference
Alexander Tschantz, Anil K. Seth, and Christopher L. Buckley · 2019
Closest in time.
Active sensing in the categorization of visual patterns
Scott Cheng-Hsin Yang, Máté Lengyel, and Daniel M Wolpert · 2019
Closest in time.
Introducing a bayesian model of selective attention based on active inference
M. Berk Mirza, Rick A. Adams, Karl Friston, and Thomas Parr · 2019
Closest in time.
Computational mechanisms of curiosity and goal-directed exploration
Philipp Schwartenbeck, Johannes Passecker, Tobias U Hauser, Thomas HB FitzGerald, Martin Kronbichler, and Karl J Friston · 2019
Closest in time.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2019
Closest in time.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H. Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, George Tucker, and Henryk Michalewski · 2019
Closest in time.
Uncertainty-aware model-based policy optimization
Tung-Long Vuong and Kenneth Tran · 2019
Closest in time.
A survey on intrinsic motivation in reinforcement learning
Arthur Aubret, Laetitia Matignon, and Salima Hassas · 2019
Closest in time.
Where does value come from?
Keno Juechems and Christopher Summerfield · 2019
Closest in time.