Fetching the paper…
Reading the bibliography…
Restless multi-armed bandits (RMAB) play a central role in modeling sequential decision making problems under an instantaneous activation constraint that at most B arms can be activated at any decision epoch.
Restless Bandits: Activity Allocation in A Changing World
Whittle, P · 1988
Earlier work this paper cites.
On An Index Policy for Restless Bandits
Weber, R. R. and Weiss, G · 1990
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W · 1994
Earlier work this paper cites.
The complexity of optimal queueing network control
Papadimitriou, C. H. and Tsitsiklis, J. N · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Convex optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2009
Earlier work this paper cites.
Empirical Bernstein Bounds and Sample Variance Penalization
Maurer, A. and Pontil, M · 2009
Earlier work this paper cites.
General notions of indexability for queueing control and asset management
Glazebrook, K. D., Hodge, D. J., and Kirkbride, C · 2011
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G., Gyorgy, A., and Szepesvári, C · 2012
Earlier work this paper cites.
Markov models for treatment adherence in obstructive sleep apnea
Kang, Y., Prabhu, V. V., Sawyer, A. M., and Griffin, P. M · 2013
Earlier work this paper cites.
Restless multi-armed bandits under time-varying activation constraints for dynamic spectrum access
Cohen, K., Zhao, Q., and Scaglione, A · 2014
Cited alongside, same era.
Index Policies for A Multi-Class Queue with Convex Holding Cost and Abandonments
Larrañaga, M., Ayesta, U., and Verloop, I. M · 2014
Cited alongside, same era.
Data-Driven Channel Modeling Using Spectrum Measurement
Sheng, S.-P., Liu, M., and Saigal, R · 2014
Cited alongside, same era.
Concentration inequalities for sums and martingales
Bercu, B., Delyon, B., Rio, E., et al · 2015
Cited alongside, same era.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Neu, G · 2015
Cited alongside, same era.
Asymptotically optimal priority policies for indexable and nonindexable restless bandits
Verloop, I. M · 2016
Online learning in mdps with linear function approximation and bandit feedback
Neu, G. and Olkhovskaya, J · 2021
Later among the works it cites.
Whittle index based q-learning for restless bandits with average reward
Avrachenkov, K. E. and Borkar, V. S · 2022
Later among the works it cites.
Near-optimal regret for adversarial mdp with delayed bandit feedback
Jin, T., Lancewicki, T., Luo, H., Mansour, Y., and Rosenberg, A · 2022
Later among the works it cites.
Towards soft fairness in restless multi-armed bandits
Li, D. and Varakantham, P · 2022
Later among the works it cites.
A best-of-both-worlds algorithm for constrained mdps with long-term constraints
Germano, J., Stradi, F. E., Genalti, G., Castiglioni, M., Marchesi, A., and Gatti, N · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning adversarial markov decision processes with bandit feedback and unknown transition
Jin, C., Jin, T., Luo, H., Sra, S., and Yu, T · 2020
Cited alongside, same era.
Upper confidence primal-dual reinforcement learning for cmdp with adversarial loss
Qiu, S., Wei, X., Yang, Z., Ye, J., and Wang, Z · 2020
Cited alongside, same era.
Restless-ucb, an efficient and low-complexity algorithm for online restless bandits
Wang, S., Huang, L., and Lui, J · 2020
Cited alongside, same era.
The best of both worlds: stochastic and adversarial episodic mdps with unknown transition
Jin, T., Huang, L., and Luo, H · 2021
Cited alongside, same era.
Beyond” To Act or Not to Act”: Fast Lagrangian Approaches to General Multi-Action Restless Bandits
Killian, J. A., Perrault, A., and Tambe, M · 2021
Cited alongside, same era.
Policy optimization in adversarial mdps: Improved exploration via dilated bonuses
Luo, H., Wei, C.-Y., and Lee, C.-W · 2021
Cited alongside, same era.
Planning to fairly allocate: Probabilistic fairness in the restless bandit setting
Herlihy, C., Prins, A., Srinivasan, A., and Dickerson, J. P · 2023
Later among the works it cites.
No-regret online reinforcement learning with adversarial losses and transitions
Jin, T., Liu, J., Rouyer, C., Chan, W., We, C.-Y., and Luo, H · 2023
Later among the works it cites.
Online resource allocation in episodic markov decision processes
Lee, D. and Lee, D · 2023
Later among the works it cites.
Markovian restless bandits and index policies: A review
Niño-Mora, J · 2023
Later among the works it cites.
Finite-time analysis of whittle index based q-learning for restless multi-armed bandits with neural network function approximation
Xiong, G. and Li, J · 2023
Later among the works it cites.
Reinforcement learning for dynamic dimensioning of cloud caches: A restless bandit approach
Xiong, G., Wang, S., Yan, G., and Li, J · 2023
Later among the works it cites.
Online restless multi-armed bandits with long-term fairness constraints
Wang, S., Xiong, G., and Li, J · 2024
Closest in time.