Fetching the paper…
Reading the bibliography…
A Robust Markov Decision Process (RMDP) is a sequential decision making model that accounts for uncertainty in the parameters of dynamic systems.
A new approach to linear filtering and prediction problems
Kalman, Rudolph Emil et al · 1960
Earlier work this paper cites.
Applied optimal estimation
Gelb, Arthur · 1974
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Zeiler, Matthew D · 1974
Earlier work this paper cites.
Optimal filtering
Anderson, Brian DO and Moore, John B · 1979
Earlier work this paper cites.
Training multilayer perceptrons with the extende kalman algorithm
Singhal, Sharad and Wu, Lance · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, Richard S · 1988
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Neuro-dynamic programming (optimization and neural computation series, 3)
Bertsekas, Dimitri P and Tsitsiklis, John N · 1996
Earlier work this paper cites.
New extension of the kalman filter to nonlinear systems
Julier, Simon J and Uhlmann, Jeffrey K · 1997
Earlier work this paper cites.
A recursive algorithm based on the extended kalman filter for the training of feedforward neural models
Rivals, Isabelle and Personnaz, Léon · 1998
Earlier work this paper cites.
Coherent measures of risk
Artzner, Philippe, Delbaen, Freddy, Eber, Jean-Marc, and Heath, David · 1999
Earlier work this paper cites.
The unscented kalman filter for nonlinear estimation
Wan, Eric A and Van Der Merwe, Rudolph · 2000
Earlier work this paper cites.
Comparison between the unscented kalman filter and the extended kalman filter for the position estimation module of an integrated navigation information system
St-Pierre, Mathieu and Gingras, Denis · 2004
Cited alongside, same era.
Sigma-point Kalman filters for probabilistic inference in dynamic state-space models
Van Der Merwe, Rudolph · 2004
Cited alongside, same era.
Robust dynamic programming
Iyengar, Garud N · 2005
Cited alongside, same era.
Robust control of markov decision processes with uncertain transition matrices
Nilim, Arnab and El Ghaoui, Laurent · 2005
Cited alongside, same era.
Robust, risk-sensitive, and data-driven control of Markov decision processes
Le Tallec, Yann · 2007
Cited alongside, same era.
Bias and variance approximation in value function estimates
Mannor, Shie, Simester, Duncan, Sun, Peng, and Tsitsiklis, John N · 2007
Markov decision processes: discrete stochastic dynamic programming
Puterman, Martin L · 2014
Later among the works it cites.
Scaling up robust mdps using function approximation
Tamar, Aviv, Mannor, Shie, Xu, Huan, and SG, EDU · 2014
Later among the works it cites.
Weight uncertainty in neural networks
Blundell, Charles, Cornebise, Julien, Kavukcuoglu, Koray, and Wierstra, Daan · 2015
Later among the works it cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Chow, Yinlam, Tamar, Aviv, Mannor, Shie, and Pavone, Marco · 2015
Later among the works it cites.
A comprehensive survey on safe reinforcement learning
Garcıa, Javier and Fernández, Fernando · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kalman temporal differences
Geist, Matthieu and Pietquin, Olivier · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Robust approximate bilinear programming for value function approximation
Petrik, Marek and Zilberstein, Shlomo · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Later among the works it cites.
Learning weight uncertainty with stochastic gradient mcmc for shape classification
Li, Chunyuan, Stevens, Andrew, Chen, Changyou, Pu, Yunchen, Gan, Zhe, and Carin, Lawrence · 2016
Later among the works it cites.
Combating reinforcement learning’s sisyphean curse with intrinsic fear
Lipton, Zachary C, Gao, Jianfeng, Li, Lihong, Chen, Jianshu, and Deng, Li · 2016
Later among the works it cites.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, Aravind, Ghotra, Sarvjeet, Levine, Sergey, and Ravindran, Balaraman · 2016
Later among the works it cites.
An overview of gradient descent optimization algorithms
Ruder, Sebastian · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, Hado, Guez, Arthur, and Silver, David · 2016
Later among the works it cites.