Fetching the paper…
Reading the bibliography…
We study the model-based undiscounted reinforcement learning for partially observable Markov decision processes (POMDPs).
Blackwell D (1965) Discounted dynamic programming. The Annals of Mathematical Statistics 36(1):226–235
1965
Earlier work this paper cites.
Azuma K (1967) Weighted sums of certain dependent random variables. Tohoku Mathematical Journal, Second Series 19(3):357–367
1967
Earlier work this paper cites.
Ross S (1968) Arbitrary state markovian decision processes. The Annals of Mathematical Statistics 39(6):2118–2122
1968
Earlier work this paper cites.
Bertsekas D (1976) Dynamic programming and stochastic control (Academic Press, New York)
1976
Earlier work this paper cites.
Zhang H, Chao X, Shi C (2020) Closing the gap: A learning algorithm for lost-sales inventory systems with lead times. Management Science 66(5):1962–1980
1980
Earlier work this paper cites.
Stephens M (2000) Dealing with label switching in mixture models. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 62(4):795–809
2000
Earlier work this paper cites.
Ormoneit D, Glynn P (2002) Kernel-based reinforcement learning in average-cost problems. IEEE Transactions on Automatic Control 47(10):1624–1636
2002
Earlier work this paper cites.
Yu H, Bertsekas D (2004) Discretized approximations for pomdp with average cost. Conference on Uncertainty in Artificial Intelligence 619–627
2004
Earlier work this paper cites.
Cappé O, Moulines E, Rydén T (2005) Inference in hidden Markov models (Springer Science & Business Media)
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
Hinderer K (2005) Lipschitz continuity of value functions in markovian decision processes. Mathematical Methods of Operations Research 62(1):3–22
2005
Earlier work this paper cites.
Auer P, Ortner R (2006) Logarithmic online regret bounds for undiscounted reinforcement learning. Advances in Neural Information Processing Systems 49–56
2006
Earlier work this paper cites.
Hsu S, Chuang D, Arapostathis A (2006) On the existence of stationary optimal policies for partially observed mdps under the long-run average cost criterion. Systems & Control Letters 55(2):165–173
2006
Earlier work this paper cites.
Cao X, Guo X (2007) Partially observable markov decision processes with reward information: Basic ideas and models. IEEE Transactions on Automatic Control 52(4):677–681
2007
Earlier work this paper cites.
Slivkins A, Upfal E (2008) Adapting to a changing environment: the brownian restless bandits. Conference on Learning Theory 343–354
2008
Earlier work this paper cites.
Yu H, Bertsekas D (2008) On near optimality of the set of finite-state controllers for average cost pomdp. Mathematics of Operations Research 33(1):1–11
2008
Earlier work this paper cites.
Guha S, Munagala K, Shi P (2010) Approximation algorithms for restless bandit problems. Journal of the ACM (JACM) 58(1):1–50
2010
Earlier work this paper cites.
Jaksch T, Ortner R, Auer P (2010) Near-optimal regret bounds for reinforcement learning. The Journal of Machine Learning Research 1563–1600
2010
Earlier work this paper cites.
Rusmevichientong P, Tsitsiklis J (2010) Linearly parameterized bandits. Mathematics of Operations Research 35(2):395–411
2010
Earlier work this paper cites.
Garivier A, Moulines E (2011) On upper-confidence bound policies for switching bandit problems. International Conference on Algorithmic Learning Theory 174–188
2011
Earlier work this paper cites.
Ross S, Pineau J, Chaib-draa B, Kreitmann P (2011) A bayesian approach for learning and planning in partially observable markov decision processes. The Journal of Machine Learning Research 1729–1770
2011
Cited alongside, same era.
Anandkumar A, Hsu D, Kakade SM (2012) A method of moments for mixture models and hidden markov models. Conference on Learning Theory 33.1––33.34
2012
Cited alongside, same era.
Bubeck S, Cesa-Bianchi N (2012) Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Machine Learning 5(1):1–122
2012
Cited alongside, same era.
Ortner R, Ryabko D (2012) Online regret bounds for undiscounted continuous reinforcement learning. Advances in Neural Information Processing Systems 1772–1780
2012
Cited alongside, same era.
Spaan M (2012) Partially observable markov decision processes. Reinforcement Learning 387–414
2012
2018
Later among the works it cites.
Igl M, Zintgraf L, Le T, Wood F, Whiteson S (2018) Deep variational reinforcement learning for pomdps. International Conference on Machine Learning 2117–2126
2018
Later among the works it cites.
Jin C, Allen-Zhu Z, Bubeck S, Jordan M (2018) Is q-learning provably efficient? Advances in Neural Information Processing Systems 4868–4878
2018
Later among the works it cites.
Sutton R, Barto A (2018) Reinforcement learning: An introduction (MIT press)
2018
Later among the works it cites.
Zhang H, Chao X, Shi C (2018) Perishable inventory systems: Convexity results for base-stock policies and learning algorithms under censored demand. Operations Research 66(5):1276–1286
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Anandkumar A, Ge R, Hsu D, Kakade SM, Telgarsky M (2014) Tensor decompositions for learning latent variable models. Journal of Machine Learning Research 2773–2832
2014
Cited alongside, same era.
Besbes O, Gur Y, Zeevi A (2014) Stochastic multi-armed-bandit problem with non-stationary rewards. Advances in Neural Information Processing Systems 199–207
2014
Cited alongside, same era.
Ortner R, Ryabko D, Auer P, Munos R (2014) Regret bounds for restless markov bandits. Theoretical Computer Science 62–76
2014
Cited alongside, same era.
Puterman M (2014) Markov decision processes: discrete stochastic dynamic programming (John Wiley & Sons)
2014
Cited alongside, same era.
Hausknecht M, Stone P (2015) Deep recurrent q-learning for partially observable mdps. Association for the Advancement of Artificial Intelligence Fall Symposium Series 29–37
2015
Cited alongside, same era.
Lakshmanan K, Ortner R, Ryabko D (2015) Improved regret bounds for undiscounted continuous reinforcement learning. International Conference on Machine Learning 524–532
2015
Cited alongside, same era.
Azizzadenesheli K, Lazaric A, Anandkumar A (2016) Reinforcement learning of pomdps using spectral methods. Conference on Learning Theory 193–256
2016
Cited alongside, same era.
2018
Later among the works it cites.
Auer P, Gajane P, Ortner R (2019) Adaptively tracking the best bandit arm with an unknown number of distribution changes. Conference on Learning Theory 138–158
2019
Later among the works it cites.
Chen B, Chao X, Ahn H (2019) Coordinating pricing and inventory replenishment with nonparametric demand learning. Operations Research 67(4):1035–1052
2019
Later among the works it cites.
Cheung W, Simchi-Levi D, Zhu R (2019) Non-stationary reinforcement learning: The blessing of (more) optimism . Available at SSRN 3397818
2019
Later among the works it cites.
Lehéricy L (2019) Consistent order estimation for nonparametric hidden markov models. Bernoulli 25(1):464–498
2019
Later among the works it cites.
Zhang Z, Ji X (2019) Regret minimization for reinforcement learning by evaluating the optimal bias function. Advances in Neural Information Processing Systems 2827–2836
2019
Later among the works it cites.
Chen W, Shi C, Duenyas I (2020) Optimal learning algorithms for stochastic inventory systems with random capacities. Production and Operations Management 29(7):1624–1649
2020
Later among the works it cites.
Jin C, Kakade S, Krishnamurthy A, Liu (2020) Sample-efficient reinforcement learning of undercomplete pomdps. Advances in Neural Information Processing Systems 18530–18539
2020
Later among the works it cites.
Lattimore T, Szepesvári C (2020) Bandit algorithms (Cambridge University Press)
2020
Later among the works it cites.
Sharma H, Jafarnia-Jahromi M, Jain R (2020) Approximate relative value learning for average-reward continuous state mdps. Conference on Uncertainty in Artificial Intelligence 956–964
2020
Later among the works it cites.
Zhu F, Zheng Z (2020) When demands evolve larger and noisier: Learning and earning in a growing environment. International Conference on Machine Learning 11629–11638
2020
Later among the works it cites.
Kwon J, Efroni Y, Caramanis C, Mannor S (2021) Rl for latent mdps: Regret guarantees and a lower bound. Advances in Neural Information Processing Systems 34
2021
Closest in time.
Nambiar M, Simchi-Levi D, Wang H (2021) Dynamic inventory allocation with demand learning for seasonal goods. Production and Operations Management 30(3):750–765
2021
Closest in time.
Zhou X, Xiong Y, Chen N, Gao X (2021) Regime switching bandits. Advances in Neural Information Processing Systems 34
2021
Closest in time.
Cheung W, Simchi-Levi D, Zhu R (2022) Hedging the drift: Learning to optimize under nonstationarity. Management Science 68(3):1696–1713
2022
Closest in time.