Fetching the paper…
Reading the bibliography…
Information-theoretic principles for learning and acting have been proposed to solve particular classes of Markov Decision Problems.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Neuro-dynamic programming
DP Bertsekas and JN Tsitsiklis · 1996
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Michael O’Gordon Duff · 2002
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Linear theory for control of nonlinear stochastic systems
Hilbert J Kappen · 2005
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2006
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Shie Mannor, Duncan Simester, Peng Sun, and John N Tsitsiklis · 2007
Earlier work this paper cites.
The many faces of optimism: a unifying approach
István Szita and András Lőrincz · 2008
Earlier work this paper cites.
Robustness
Lars Peter Hansen and Thomas J Sargent · 2008
Earlier work this paper cites.
Efficient computation of optimal actions
Emanuel Todorov · 2009
Earlier work this paper cites.
Reinforcement learning in finite mdps: Pac analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Cited alongside, same era.
Risk sensitive path integral control
Bart van den Broek, Wim Wiegerinck, and Hilbert J. Kappen · 2010
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
István Szita and Csaba Szepesvári · 2010
Cited alongside, same era.
Relative entropy policy search
J Peters, K Mülling, Y Altun, Fox D Poole, et al · 2010
Cited alongside, same era.
A minimum relative entropy principle for learning and acting
Pedro A Ortega and Daniel A Braun · 2010
Cited alongside, same era.
A bayesian rule for adaptive control based on causal interventions
Pedro A Ortega and Daniel A Braun · 2010
Cited alongside, same era.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Later among the works it cites.
Thermodynamics as a theory of decision-making with information-processing costs
Pedro A Ortega and Daniel A Braun · 2013
Later among the works it cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Later among the works it cites.
Adaptive control
Karl J Åström and Björn Wittenmark · 2013
Later among the works it cites.
Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search
Arthur Guez, David Silver, and Peter Dayan · 2013
Later among the works it cites.
Generalized thompson sampling for sequential decision-making and causal inference
Pedro A Ortega and Daniel A Braun · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Path integral control and bounded rationality
Daniel A Braun, Pedro A Ortega, Evangelos Theodorou, and Stefan Schaal · 2011
Cited alongside, same era.
Information theory of decisions and actions
Naftali Tishby and Daniel Polani · 2011
Cited alongside, same era.
A bayesian approach for learning and planning in partially observable markov decision processes
Stéphane Ross, Joelle Pineau, Brahim Chaib-draa, and Pierre Kreitmann · 2011
Cited alongside, same era.
Trading value and information in mdps
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Cited alongside, same era.
Robustness and risk-sensitivity in markov decision processes
Takayuki Osogami · 2012
Cited alongside, same era.
Risk-sensitive reinforcement learning
Yun Shen, Michael J Tobia, Tobias Sommer, and Klaus Obermayer · 2014
Later among the works it cites.
Monte carlo methods for exact & efficient solution of the generalized optimality equations
Pedro A Ortega, Daniel A Braun, and Naftali Tishby · 2014
Later among the works it cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Later among the works it cites.
G-learning: Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Later among the works it cites.
Rlpy: A value-function-based reinforcement learning framework for education and research
Alborz Geramifard, Christoph Dann, Robert H Klein, William Dabney, and Jonathan P How · 2015
Later among the works it cites.