Fetching the paper…
Reading the bibliography…
Large-scale Markov decision processes (MDPs) require planning algorithms with runtime independent of the number of states of the MDP.
Large-scale Markov decision problems via the linear programming dual (Jan. 2019)
Abbasi-Yadkori, Y., Bartlett, P. L., Chen, X., and Malek, A · 1901
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation (2019)
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 1907
Earlier work this paper cites.
Faster saddle-point optimization for solving large-scale Markov decision processes (Jan. 2020)
Bas-Serrano, J. and Neu, G · 1909
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning? (2019a)
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F · 1910
Earlier work this paper cites.
Learning with good feature representations in bandits and in RL with a generative model
Lattimore, T., Szepesvári, Cs., and Weisz, G · 1911
Earlier work this paper cites.
Comments on the Du-Kakade-Wang-Yang lower bounds (2019)
Van Roy, B. and Dong, S · 1911
Earlier work this paper cites.
Direct value-approximation for factored MDPs
Schuurmans, D. and Patrascu, R · 1981
Earlier work this paper cites.
Generalized polynomial approximations in Markovian decision processes
Schweitzer, P. J. and Seidmann, A · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
A survey of applications of Markov decision processes
White, D. J · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Singh, S. P., Jaakkola, T., and Jordan, M. I · 1995
Earlier work this paper cites.
Numerical dynamic programming in economics
Rust, J · 1996
Earlier work this paper cites.
A survey of computational complexity results in systems and control
Blondel, V. D. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Efficient approximate planning in continuous space Markovian decision problems
Szepesvári, Cs · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Cited alongside, same era.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Kearns, M., Mansour, Y., and Ng, A. Y · 2002
Cited alongside, same era.
The linear programming approach to approximate dynamic programming
de Farias, D. P. and Van Roy, B · 2003
Cited alongside, same era.
Efficient solution algorithms for factored MDPs
Guestrin, C., Koller, D., Parr, R., and Venkataraman, S · 2003
Cited alongside, same era.
On constraint sampling in the linear programming approach to approximate dynamic programming
de Farias, D. P. and Van Roy, B · 2004
Cited alongside, same era.
Bandit based Monte-Carlo planning
Kocsis, L. and Szepesvári, Cs · 2006
Cited alongside, same era.
The strong convexity of von Neumann’s entropy (Jun. 2013)
Yu, Y.-L · 2013
Later among the works it cites.
Linear programming for large-scale Markov decision problems
Abbasi-Yadkori, Y., Bartlett, P. L., and Malek, A · 2014
Later among the works it cites.
Simple regret optimization in online planning for Markov decision processes
Feldman, Z. and Domshlak, C · 2014
Later among the works it cites.
From bandits to Monte-Carlo tree search: The optimistic principle applied to optimization and planning
Munos, R · 2014
Later among the works it cites.
Primal-dual algorithms for discounted Markov decision processes
Cogill, R · 2015
Later among the works it cites.
Stochastic primal-dual methods and sample complexity of reinforcement learning (Dec. 2016)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stable dual dynamic programming
Wang, T., Bowling, M., Schuurmans, D., and Lizotte, D. J · 2008
Cited alongside, same era.
A smoothed approximate linear program
Desai, V. V., Farias, V. F., and Moallemi, C. C · 2009
Cited alongside, same era.
Constraint relaxation in approximate linear programs
Petrik, M. and Zilberstein, S · 2009
Cited alongside, same era.
Solving variational inequalities with Stochastic Mirror-Prox algorithm
Juditsky, A., Nemirovski, A., and Tauvel, C · 2011
Cited alongside, same era.
Non-parametric approximate dynamic programming via the kernel method
Bhat, N., Farias, V., and Moallemi, C. C · 2012
Cited alongside, same era.
Planning with Markov Decision Processes: An AI Perspective , vol. 17 of Synthesis Lectures on Artificial Intelligence and Machine Learning
Mausam and Kolobov, A · 2012
Cited alongside, same era.
Chen, Y. and Wang, M · 2016
Later among the works it cites.
Markov Decision Processes in Practice , vol. 248 of International Series in Operations Research & Management Science
Boucherie, R. J. and van Dijk, N. M., eds · 2017
Later among the works it cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Later among the works it cites.
Markov decision processes (2017)
Kallenberg, L · 2017
Later among the works it cites.
A linearly relaxed approximate linear program for Markov decision processes
Lakshminarayanan, C., Bhatnagar, S., and Szepesvári, Cs · 2017
Later among the works it cites.
Scalable bilinear π \pi learning using state and action features
Chen, Y., Li, L., and Wang, M · 2018
Later among the works it cites.
Optimizing over a restricted policy class in MDPs
Banijamali, E., Abbasi-Yadkori, Y., Ghavamzadeh, M., and Vlassis, N · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Later among the works it cites.
Limiting extrapolation in linear approximate value iteration
Zanette, A., Lazaric, A., Kochenderfer, M. J., and Brunskill, E · 2019
Later among the works it cites.