Fetching the paper…
Reading the bibliography…
Models of many real-life applications, such as queuing models of communication networks or computing systems, have a countably infinite state-space.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Dynamic programming for countable state systems
Ashok P. Maitra · 1963
Earlier work this paper cites.
An example in denumerable decision processes
Lloyd Fisher and Sheldon M Ross · 1968
Earlier work this paper cites.
Applying a new device in the optimization of exponential queuing systems
Steven A Lippman · 1975
Earlier work this paper cites.
A simple dynamic routing problem
Anthony Ephremides, Pravin Varaiya, and Jean Walrand · 1980
Earlier work this paper cites.
Control of multiple exponential servers with application to computer systems
Ronald Larsen · 1981
Earlier work this paper cites.
Hitting-time and occupation-time bounds implied by drift analysis with applications
Bruce Hajek · 1982
Earlier work this paper cites.
A new family of optimal adaptive controllers for Markov chains
P R Kumar and A Becker · 1982
Earlier work this paper cites.
Optimal adaptive controllers for unknown Markov chains
P R Kumar and Woei Lin · 1982
Earlier work this paper cites.
Optimal control of a queueing system with two heterogeneous servers
Woei Lin and P R Kumar · 1984
Earlier work this paper cites.
Asymptotically efficient adaptive allocation schemes for controlled Markov chains: Finite parameter space
R. Agrawal, D. Teneketzis, and V. Anantharam · 1989
Earlier work this paper cites.
Certainty equivalence control with forcing: Revisited
Rajeev Agrawal and Demosthenis Teneketzis · 1989
Earlier work this paper cites.
Necessary conditions for the optimality equation in average-reward Markov decision processes
Rolando Cavazos-Cadena · 1989
Earlier work this paper cites.
Weak conditions for the existence of optimal stationary policies in average Markov decision chains with unbounded costs
Rolando Cavazos-Cadena · 1989
Earlier work this paper cites.
Average cost optimal stationary policies in infinite state Markov decision processes with unbounded costs
Linn I Sennott · 1989
Earlier work this paper cites.
The Kumar-Becker-Lin scheme revisited
V S Borkar · 1990
Earlier work this paper cites.
Yet another application of a binomial recurrence order statistics
Wojciech Szpankowski and Vernon Rego · 1990
Earlier work this paper cites.
Comparing recent assumptions for the existence of average optimal stationary policies
Rolando Cavazos-Cadena and Linn I Sennott · 1992
Earlier work this paper cites.
On ergodicity and recurrence properties of a Markov chain by an application to an open Jackson network
Arie Hordijk and Flora Spieksma · 1992
Earlier work this paper cites.
Jointly optimal routing and scheduling in packet ratio networks
Leandros Tassiulas and Anthony Ephremides · 1992
Earlier work this paper cites.
Discrete-time controlled Markov processes with average cost criterion: A survey
Aristotle Arapostathis, Vivek S Borkar, Emmanuel Fernández-Gaucherand, Mrinal K Ghosh, and Steven I Marcus · 1993
Earlier work this paper cites.
A survey of Markov decision models for control of networks of queues
Shaler Stidham and Richard Weber · 1993
Earlier work this paper cites.
Dynamic server allocation to parallel queues with randomly varying connectivity
Leandros Tassiulas and Anthony Ephremides · 1993
Cited alongside, same era.
Machine learning and nonparametric bandit theory
Tze-Leung Lai and Sidney Yakowitz · 1995
Cited alongside, same era.
Asymptotically efficient adaptive choice of control laws incontrolled Markov chains
Todd L Graves and Tze-Leung Lai · 1997
Cited alongside, same era.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Cited alongside, same era.
Polynomial convergence rates of Markov chains
Soren F Jarner and Gareth O Roberts · 2002
Cited alongside, same era.
The Poisson equation for countable Markov chains: Probabilistic methods and interpretations
Armand M Makowski and Adam Shwartz · 2002
Cited alongside, same era.
On learning the c μ \mu rule in single and parallel server networks
Subhashini Krishnasamy, Ari Arapostathis, Ramesh Johari, and Sanjay Shakkottai · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
Scalar posterior sampling with applications
Georgios Theocharous, Zheng Wen, Yasin Abbasi Yadkori, and Nikos Vlassis · 2018
Later among the works it cites.
Learning in structured MDPs with convex cost functions: Improved regret bounds for inventory management
Shipra Agrawal and Randy Jia · 2019
Later among the works it cites.
Stable reinforcement learning with unbounded state space
Devavrat Shah, Qiaomin Xie, and Zhi Xu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inequalities for the L 1 L_{1} deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Cited alongside, same era.
A minimum relative entropy principle for learning and acting
Pedro A Ortega and Daniel A Braun · 2010
Cited alongside, same era.
Markov chains and stochastic stability
Sean P Meyn and Richard L Tweedie · 2012
Cited alongside, same era.
Max-Weight learning algorithms for scheduling in unknown environments
Michael J Neely, Scott T Rager, and Thomas F La Porta · 2012
Cited alongside, same era.
(More) Efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Job dispatching policies for queueing systems with unknown service rates
Tuhinangshu Choudhury, Gauri Joshi, Weina Wang, and Sanjay Shakkottai · 2021
Later among the works it cites.
Reinforcement learning in parametric MDPs with exponential families
Sayak Ray Chowdhury, Aditya Gopalan, and Odalric-Ambrym Maillard · 2021
Later among the works it cites.
Learning unknown service rates in queues: A multiarmed bandit approach
Subhashini Krishnasamy, Rajat Sen, Ramesh Johari, and Sanjay Shakkottai · 2021
Later among the works it cites.
Reward biased maximum likelihood estimation for reinforcement learning
Akshay Mete, Rahul Singh, Xi Liu, and P R Kumar · 2021
Later among the works it cites.
Learning algorithms for minimizing queue length regret
Thomas Stahlbuhk, Brooke Shrader, and Eytan Modiano · 2021
Later among the works it cites.
Learning and information in stochastic networks and queues
Neil Walton and Kuang Xu · 2021
Later among the works it cites.
Learning a discrete set of optimal allocation rules in queueing systems with unknown service rates
Saghar Adler, Mehrdad Moharrami, and Vijay Subramanian · 2022
Later among the works it cites.
On learning Whittle index policy for restless bandits with scalable regret
Nima Akbarzadeh and Aditya Mahajan · 2022
Later among the works it cites.
Queueing network controls via deep reinforcement learning
Jim G Dai and Mark Gluzman · 2022
Later among the works it cites.
Efficient decentralized multi-agent learning in asymmetric queuing systems
Daniel Freund, Thodoris Lykouris, and Wentao Weng · 2022
Later among the works it cites.
Online learning for unknown partially observable MDPs
Mehdi Jafarnia Jahromi, Rahul Jain, and Ashutosh Nayyar · 2022
Later among the works it cites.
Bilinear exponential family of MDPs: Frequentist regret bound with tractable exploration and planning
Reda Ouhamma, Debabrota Basu, and Odalric Maillard · 2023
Closest in time.
Efficient online learning with offline datasets for infinite horizon MDPs: A Bayesian approach
Dengwang Tang, Rahul Jain, Botao Hao, and Zheng Wen · 2023
Closest in time.
Learning while scheduling in multi-server systems with unknown statistics: Max-Weight with discounted UCB
Zixian Yang, R Srikant, and Lei Ying · 2023
Closest in time.
Learning-based optimal admission control in a single-server queuing system
Asaf Cohen, Vijay Subramanian, and Yili Zhang · 2024
Closest in time.