Fetching the paper…
Reading the bibliography…
The restless multi-armed bandit (RMAB) framework is a popular model with applications across a wide variety of fields.
Restless Bandits: Activity Allocation in a Changing World
Whittle, P. 1988 · 1988
Earlier work this paper cites.
On an index policy for restless bandits
Weber, R. R.; and Weiss, G. 1990 · 1990
Earlier work this paper cites.
The Complexity of Optimal Queuing Network Control
Papadimitriou, C. H.; and Tsitsiklis, J. N. 1999 · 1999
Earlier work this paper cites.
Learning algorithms for Markov decision processes with average cost
Abounadi, J.; Bertsekas, D.; and Borkar, V. S. 2001 · 2001
Earlier work this paper cites.
Restless bandits, partial conservation laws and indexability
Niño-Mora, J. 2001 · 2001
Earlier work this paper cites.
Some Indexable Families of Restless Bandit Problems
Glazebrook, K. D.; Ruiz-Hernandez, D.; and Kirkbride, C. 2006 · 2006
Earlier work this paper cites.
Dynamic priority allocation via restless bandit marginal productivity indices
Niño-Mora, J. 2007 · 2007
Earlier work this paper cites.
Relaxations of weakly coupled stochastic dynamic programs
Adelman, D.; and Mersereau, A. J. 2008 · 2008
Earlier work this paper cites.
Indexability of restless bandit problems and optimality of whittle index for dynamic multichannel access
Liu, K.; and Zhao, Q. 2010 · 2010
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Indexability and optimal index policies for a class of reinitialising restless bandits
Villar, S. S. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning for building HVAC control
Wei, T.; Wang, Y.; and Zhu, Q. 2017 · 2017
Cited alongside, same era.
DeepCAS: A Deep Reinforcement Learning Algorithm for Control-Aware Scheduling
Demirel, B.; Ramaswamy, A.; Quevedo, D. E.; and Karl, H. 2018 · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Cited alongside, same era.
Deep reinforcement learning for dynamic multichannel access in wireless networks
Wang, S.; Liu, H.; Gomes, P. H.; and Krishnamachari, B. 2018 · 2018
Cited alongside, same era.
Towards Q-learning the Whittle Index for Restless Bandits
Fu, J.; Nazarathy, Y.; Moka, S.; and Taylor, P. G. 2019 · 2019
Cited alongside, same era.
A whittle index approach to minimizing functions of age of information
Tripathi, V.; and Modiano, E. 2019 · 2019
Dual-mandate patrols: Multi-armed bandits for green security
Xu, L.; Bondi, E.; Fang, F.; Perrault, A.; Wang, K.; and Tambe, M. 2021 · 2021
Later among the works it cites.
Whittle index based Q-learning for restless bandits with average reward
Avrachenkov, K. E.; and Borkar, V. S. 2022 · 2022
Later among the works it cites.
Uncertainty-of-Information Scheduling: A Restless Multiarmed Bandit Framework
Chen, G.; Liew, S. C.; and Shao, Y. 2022 · 2022
Later among the works it cites.
Field study in deploying restless multi-armed bandits: Assisting non-profits in improving maternal and child health
Mate, A.; Madaan, L.; Taneja, A.; Madhiwalla, N.; Verma, S.; Singh, G.; Hegde, A.; Varakantham, P.; and Tambe, M. 2022 · 2022
Later among the works it cites.
Reinforcement learning augmented asymptotically optimal index policy for finite-horizon restless bandits
Xiong, G.; Li, J.; and Singh, R. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Whittle index policy for dynamic multichannel allocation in remote state estimation
Wang, J.; Ren, X.; Mo, Y.; and Shi, L. 2019 · 2019
Cited alongside, same era.
Deep reinforcement learning for wireless sensor scheduling in cyber–physical systems
Leong, A. S.; Ramaswamy, A.; Quevedo, D. E.; Karl, H.; and Shi, L. 2020 · 2020
Cited alongside, same era.
Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare
Biswas, A.; Aggarwal, G.; Varakantham, P.; and Tambe, M. 2021 · 2021
Cited alongside, same era.
NeurWIN: Neural Whittle index network for restless bandits via deep RL
Nakhleh, K.; Ganji, S.; Hsieh, P.-C.; Hou, I.; Shakkottai, S.; et al. 2021 · 2021
Cited alongside, same era.
Restless Multi-Armed Bandit in Opportunistic Scheduling
Wang, K.; and Chen, L. 2021 · 2021
Cited alongside, same era.
Index-aware reinforcement learning for adaptive video streaming at the wireless edge
Xiong, G.; Qin, X.; Li, B.; Singh, R.; and Li, J. 2022 · 2022
Later among the works it cites.
Learning infinite-horizon average-reward restless multi-action bandits via index awareness
Xiong, G.; Wang, S.; and Li, J. 2022 · 2022
Later among the works it cites.
Optimistic whittle index policy: Online learning for restless bandits
Wang, K.; Xu, L.; Taneja, A.; and Tambe, M. 2023 · 2023
Later among the works it cites.
Finite-time analysis of whittle index based Q-learning for restless multi-armed bandits with neural network function approximation
Xiong, G.; and Li, J. 2023 · 2023
Later among the works it cites.
An Index Policy for Minimizing the Uncertainty-of-Information of Markov Sources
Chen, G.; and Liew, S. C. 2024 · 2024
Closest in time.
Whittle Index-Based Q-Learning for Wireless Edge Caching With Linear Function Approximation
Xiong, G.; Wang, S.; Li, J.; and Singh, R. 2024 · 2024
Closest in time.