Fetching the paper…
Reading the bibliography…
On the theory of dynamic programming
Richard Bellman · 1952
Earlier work this paper cites.
Optimal control of markov processes with incomplete state information
Karl J Astrom · 1965
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Attractors for random dynamical systems
Hans Crauel and Franco Flandoli · 1994
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Efficient model-based exploration
Marco Wiering and Jürgen Schmidhuber · 1998
Earlier work this paper cites.
Bayesian q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
Planning by probabilistic inference
Hagai Attias · 2003
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Matthew J Beal · 2003
Earlier work this paper cites.
Shaping and policy search in reinforcement learning
Andrew Y Ng and Michael I Jordan · 2003
Earlier work this paper cites.
Whence the expected free energy?
Beren Millidge, Alexander Tschantz, and Christopher L Buckley · 2004
Earlier work this paper cites.
Multi-armed bandit algorithms and empirical evaluation
Joannes Vermorel and Mehryar Mohri · 2005
Earlier work this paper cites.
On the relationship between active inference and control as inference
Beren Millidge, Alexander Tschantz, Anil K Seth, and Christopher L Buckley · 2006
Earlier work this paper cites.
Representation and timing in theories of the dopamine system
Nathaniel D Daw, Aaron C Courville, and David S Touretzky · 2006
Earlier work this paper cites.
Developmental robotics, optimal artificial curiosity, creativity, music, and the fine arts
Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Context learning in the rodent hippocampus
Mark C Fuhs and David S Touretzky · 2007
Earlier work this paper cites.
Bayes-adaptive pomdps
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Earlier work this paper cites.
Markov decision processes: a tool for sequential decision making under uncertainty
Oguzhan Alagoz, Heather Hsu, Andrew J Schaefer, and Mark S Roberts · 2010
Earlier work this paper cites.
Learning latent structure: carving nature at its joints
Samuel J Gershman and Yael Niv · 2010
Earlier work this paper cites.
Action understanding and active inference
Karl Friston, Jérémie Mattout, and James Kilner · 2011
Earlier work this paper cites.
Model-based influences on humans’ choices and striatal prediction errors
Nathaniel D Daw, Samuel J Gershman, Ben Seymour, Peter Dayan, and Raymond J Dolan · 2011
Earlier work this paper cites.
Planning as inference
Matthew Botvinick and Marc Toussaint · 2012
Earlier work this paper cites.
Stochastic thermodynamics, fluctuation theorems and molecular machines
Udo Seifert · 2012
Earlier work this paper cites.
Bayesian hierarchical reinforcement learning
Feng Cao and Soumya Ray · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Active inference and agency: optimal control without cost functions
Karl Friston, Spyridon Samothrakis, and Read Montague · 2012
Cited alongside, same era.
Complex inference in neural circuits with probabilistic population codes and topic models
Jeff Beck, Alexandre Pouget, and Katherine A Heller · 2012
Cited alongside, same era.
Variance-based rewards for approximate bayesian reinforcement learning
Jonathan Sorg, Satinder Singh, and Richard L Lewis · 2012
Cited alongside, same era.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Later among the works it cites.
The free energy principle for action and perception: A mathematical review
Christopher L Buckley, Chang Sub Kim, Simon McGregor, and Anil K Seth · 2017
Later among the works it cites.
Reinforcement learning and episodic memory in humans and animals: an integrative framework
Samuel J Gershman and Nathaniel D Daw · 2017
Later among the works it cites.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Context-dependent decision-making: a simple bayesian model
Kevin Lloyd and David S Leslie · 2013
Cited alongside, same era.
Model predictive control
Eduardo F Camacho and Carlos Bordons Alba · 2013
Cited alongside, same era.
Model-based bayesian exploration
Richard Dearden, Nir Friedman, and David Andre · 2013
Cited alongside, same era.
The anatomy of choice: dopamine and decision-making
Karl Friston, Philipp Schwartenbeck, Thomas FitzGerald, Michael Moutoussis, Timothy Behrens, and Raymond J Dolan · 2014
Cited alongside, same era.
Modeling human plan recognition using bayesian theory of mind
Chris L Baker and Joshua B Tenenbaum · 2014
Cited alongside, same era.
A formal model of interpersonal inference
Michael Moutoussis, Nelson Jesús Trujillo-Barreto, Wael El-Deredy, Raymond Dolan, and Karl Friston · 2014
Cited alongside, same era.
Active inference and epistemic value
Karl Friston, Francesco Rigoli, Dimitri Ognibene, Christoph Mathys, Thomas Fitzgerald, and Giovanni Pezzulo · 2015
Cited alongside, same era.
Later among the works it cites.
Complex probabilistic inference
Samuel J Gershman and Jeffrey M Beck · 2017
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Deep variational reinforcement learning for pomdps
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Lecture slides on bayesian reinforcement learning from cs885:
Pascal Poupart · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Deep active inference
Kai Ueltzhöffer · 2018
Later among the works it cites.
The discrete and continuous brain: from decisions to movement—and back again
Thomas Parr and Karl J Friston · 2018
Later among the works it cites.
Active inference in openai gym: a paradigm for computational investigations into psychiatric illness
Maell Cullen, Ben Davey, Karl J Friston, and Rosalyn J Moran · 2018
Later among the works it cites.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2018
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
A free energy principle for a particular physics
Karl Friston · 2019
Closest in time.
Computational mechanisms of curiosity and goal-directed exploration
Philipp Schwartenbeck, Johannes Passecker, Tobias U Hauser, Thomas HB FitzGerald, Martin Kronbichler, and Karl J Friston · 2019
Closest in time.
Reinforcement learning in non-stationary environments
Sindhu Padakandla, Shalabh Bhatnagar, et al · 2019
Closest in time.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Closest in time.
Bayesian curiosity for efficient exploration in reinforcement learning
Tom Blau, Lionel Ott, and Fabio Ramos · 2019
Closest in time.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2019
Closest in time.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Closest in time.
Neuronal message passing using mean-field, bethe, and marginal approximations
Thomas Parr, Dimitrije Markovic, Stefan J Kiebel, and Karl J Friston · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Closest in time.
Making sense of reinforcement learning and probabilistic inference
Brendan O’Donoghue, Ian Osband, and Catalin Ionescu · 2020
Closest in time.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Closest in time.
Active inference on discrete state-spaces: a synthesis
Lancelot Da Costa, Thomas Parr, Noor Sajid, Sebastijan Veselic, Victorita Neacsu, and Karl Friston · 2020
Closest in time.