Fetching the paper…
Reading the bibliography…
We study reinforcement learning in non-episodic factored Markov decision processes (FMDPs).
A model for reasoning about persistence and causation
Dean, T. and Kanazawa, K. (1989) · 1989
Earlier work this paper cites.
Learning dynamic bayesian networks
Ghahramani, Z. (1997) · 1997
Earlier work this paper cites.
The complexity of plan existence and evaluation in robabilistic domains
Goldsmith, J., Littman, M. L., and Mundhenk, M. (1997) · 1997
Earlier work this paper cites.
Probabilistic propositional planning: representations and complexity
Littman, M. L. (1997) · 1997
Earlier work this paper cites.
Efficient reinforcement learning in factored mdps
Kearns, M. and Koller, D. (1999) · 1999
Earlier work this paper cites.
Stochastic dynamic programming with factored representations
Boutilier, C., Dearden, R., and Goldszmidt, M. (2000) · 2000
Earlier work this paper cites.
Bounded-parameter markov decision processes
Givan, R., Leach, S., and Dean, T. (2000) · 2000
Earlier work this paper cites.
Max-norm projections for factored mdps
Guestrin, C., Koller, D., and Parr, R. (2001) · 2001
Earlier work this paper cites.
A probabilistic analysis of bias optimality in unichain markov decision processes
Lewis, M. E. and Puterman, M. L. (2001) · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Direct value-approximation for factored mdps
Schuurmans, D. and Patrascu, R. (2002) · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored mdps
Guestrin, C., Koller, D., Parr, R., and Venkataraman, S. (2003) · 2003
Cited alongside, same era.
Model-based reinforcement learning in factored-state mdps
Strehl, A. L. (2007) · 2007
Cited alongside, same era.
Efficient structure learning in factored-state mdps
Strehl, A. L., Diuk, C., and Littman, M. L. (2007) · 2007
Cited alongside, same era.
Multi-armed bandits in metric spaces
Kleinberg, R., Slivkins, A., and Upfal, E. (2008) · 2008
Cited alongside, same era.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, P. L. and Tewari, A. (2009) · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Factored mdps for detecting topics of user sessions
Tavakol, M. and Brefeld, U. (2014) · 2014
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E. (2015) · 2015
Later among the works it cites.
Off-policy model-based learning under unknown factored dynamics
Hallak, A., Schnitzler, F., Mann, T., and Mannor, S. (2015) · 2015
Later among the works it cites.
Improved regret bounds for oracle-based adversarial contextual bandits
Syrgkanis, V., Luo, H., Krishnamurthy, A., and Schapire, R. E. (2016) · 2016
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R. (2017) · 2017
Later among the works it cites.
Learning unknown markov decision processes: A thompson sampling approach
Ouyang, Y., Gagrani, M., Nayyar, A., and Jain, R. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structure learning in ergodic factored mdps without knowledge of the transition function’s in-degree
Chakraborty, D. and Stone, P. (2011) · 2011
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B. (2013) · 2013
Cited alongside, same era.
Near-optimal reinforcement learning in factored mdps
Osband, I. and Van Roy, B. (2014) · 2014
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Multiagent planning with factored mdps
Guestrin, C., Koller, D., and Parr, R. (2002a)
Cited in the paper.
Algorithm-directed exploration for model-based reinforcement learning in factored mdps
Guestrin, C., Patrascu, R., and Schuurmans, D. (2002b)
Cited in the paper.
Later among the works it cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R., Pirotta, M., Lazaric, A., and Ortner, R. (2018) · 2018
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
Luo, H., Wei, C.-Y., Agarwal, A., and Langford, J. (2018) · 2018
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J. (2019) · 2019
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z. and Ji, X. (2019) · 2019
Later among the works it cites.