Fetching the paper…
Reading the bibliography…
We propose a new method to study the internal memory used by reinforcement learning policies.
Learning without state-estimation in partially observable Markovian decision processes
S.P. Singh, T. Jaakkola, and M.I. Jordan · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming , volume 10
Martin L Puterman · 1995
Earlier work this paper cites.
The Sciences of The Artifical
Herbert Simon · 1996
Earlier work this paper cites.
Computational mechanics: Pattern and prediction, structure and simplicity
Cosma Rohilla Shalizi and James P. Crutchfield · 2001
Earlier work this paper cites.
Entropy and inference, revisited
Ilya Nemenman, Fariel Shafee, and William Bialek · 2002
Earlier work this paper cites.
Entropy Estimates from Insufficient Samplings
Peter Grassberger · 2003
Earlier work this paper cites.
Variational information maximization for neural coding
Felix Agakov and David Barber · 2004
Cited alongside, same era.
Entropy inference and the James-Stein estimator, with application to nonlinear gene association networks
Jean Hausser and Korbinian Strimmer · 2009
Cited alongside, same era.
Qualitative Analysis of Partially-Observable Markov Decision Processes
Krishnendu Chatterjee, Laurent Doyen, and Thomas A Henzinger · 2010
Cited alongside, same era.
Information Theory of Decisions and Actions
Naftali Tishby and Daniel Polani · 2011
Cited alongside, same era.
Improved information gain estimates for decision tree induction
Sebastian Nowozin · 2012
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei a Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Later among the works it cites.
Recurrent Reinforcement Learning: A Hybrid Approach
Xiujun Li, Lihong Li, Jianfeng Gao, Xiaodong He, Jianshu Chen, Li Deng, and Ji He · 2015
Later among the works it cites.
Contextual-MDPs for PAC-Reinforcement Learning with Rich Observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Closest in time.
Markov chain order estimation with conditional mutual information
M. Papapetrou and D. Kugiumtzis · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…