Fetching the paper…
Reading the bibliography…
This paper studies restless multi-armed bandit (RMAB) problems with unknown arm transition dynamics but with known correlated arm features.
Bai, S.; Kolter, J. Z.; and Koltun, V. 2019 · 1909
Earlier work this paper cites.
Differentiable convex optimization layers
Agrawal, A.; Amos, B.; Barratt, S.; Boyd, S.; Diamond, S.; and Kolter, Z. 2019 · 1910
Earlier work this paper cites.
Learning in a changing world: Restless multiarmed bandit with unknown dynamics
Liu, H.; Liu, K.; and Zhao, Q. 2012 · 1916
Earlier work this paper cites.
State of the art—a survey of partially observable Markov decision processes: theory, models, and algorithms
Monahan, G. E. 1982 · 1982
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Whittle, P. 1988 · 1988
Earlier work this paper cites.
On an index policy for restless bandits
Weber, R. R.; and Weiss, G. 1990 · 1990
Earlier work this paper cites.
The complexity of optimal queueing network control
Papadimitriou, C. H.; and Tsitsiklis, J. N. 1994 · 1994
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S.; Barto, A. G.; et al. 1998 · 1998
Earlier work this paper cites.
Popcorn: Partially observed prediction constrained reinforcement learning
Futoma, J.; Hughes, M. C.; and Doshi-Velez, F. 2020 · 2001
Earlier work this paper cites.
Objective mismatch in model-based reinforcement learning
Lambert, N.; Amos, B.; Yadan, O.; and Calandra, R. 2020 · 2002
Earlier work this paper cites.
Differentiable top-k operator with optimal transport
Xie, Y.; Dai, H.; Chen, M.; Dai, B.; Zhao, T.; Zha, H.; Wei, W.; and Pfister, T. 2020 · 2002
Earlier work this paper cites.
Faster dynamic matrix inverse for faster lps
Jiang, S.; Song, Z.; Weinstein, O.; and Zhang, H. 2020 · 2004
Earlier work this paper cites.
Some indexable families of restless bandit problems
Glazebrook, K. D.; Ruiz-Hernandez, D.; and Kirkbride, C. 2006 · 2006
Earlier work this paper cites.
Indexability of restless bandit problems and optimality of whittle index for dynamic multichannel access
Liu, K.; and Zhao, Q. 2010 · 2010
Earlier work this paper cites.
The non-Bayesian restless multi-armed bandit: A case of near-logarithmic regret
Dai, W.; Gai, Y.; Krishnamachari, B.; and Zhao, Q. 2011 · 2011
Earlier work this paper cites.
Multi-armed bandit allocation indices
Gittins, J.; Glazebrook, K.; and Weber, R. 2011 · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S.; and Cesa-Bianchi, N. 2012 · 2012
Cited alongside, same era.
Online learning of rested and restless bandits
Tekin, C.; and Liu, M. 2012 · 2012
Cited alongside, same era.
The restless multi-armed bandit formulation of the cognitive compressive sensing problem
Bagheri, S.; and Scaglione, A. 2015 · 2015
Cited alongside, same era.
Iterative Bregman projections for regularized transportation problems
Benamou, J.-D.; Carlier, G.; Cuturi, M.; Nenna, L.; and Peyré, G. 2015 · 2015
Cited alongside, same era.
An order optimal policy for exploiting idle spectrum in cognitive radio networks
Addressing the loss-metric mismatch with adaptive loss alignment
Huang, C.; Zhai, S.; Talbott, W.; Martin, M. B.; Sun, S.-Y.; Guestrin, C.; and Susskind, J. 2019 · 2019
Later among the works it cites.
Survey on deep learning with class imbalance
Johnson, J. M.; and Khoshgoftaar, T. M. 2019 · 2019
Later among the works it cites.
Opportunistic scheduling revisited using restless bandits: Indexability and index policy
Wang, K.; Yu, J.; Chen, L.; Zhou, P.; Ge, X.; and Win, M. Z. 2019 · 2019
Later among the works it cites.
Melding the data-decisions pipeline: Decision-focused learning for combinatorial optimization
Wilder, B.; Dilkina, B.; and Tambe, M. 2019 · 2019
Later among the works it cites.
Decision trees for decision-making under the predict-then-optimize framework
Elmachtoub, A.; Liang, J. C. N.; and McNellis, R. 2020 · 2020
Later among the works it cites.
Smart predict-and-optimize for hard combinatorial optimization problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oksanen, J.; and Koivunen, V. 2015 · 2015
Cited alongside, same era.
Nonparametric bayesian reward segmentation for skill discovery using inverse reinforcement learning
Ranchod, P.; Rosman, B.; and Konidaris, G. 2015 · 2015
Cited alongside, same era.
Safe reinforcement learning
Thomas, P. S. 2015 · 2015
Cited alongside, same era.
Restless poachers: Handling exploration-exploitation tradeoffs in security domains
Qian, Y.; Zhang, C.; Krishnamachari, B.; and Tambe, M. 2016 · 2016
Cited alongside, same era.
Optnet: Differentiable optimization as a layer in neural networks
Amos, B.; and Kolter, J. Z. 2017 · 2017
Cited alongside, same era.
Task-based end-to-end model learning in stochastic optimization
Donti, P. L.; Amos, B.; and Kolter, J. Z. 2017 · 2017
Cited alongside, same era.
Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning
Krishnan, S.; Garg, A.; Patil, S.; Lea, C.; Hager, G.; Abbeel, P.; and Goldberg, K. 2017 · 2017
Cited alongside, same era.
Mandi, J.; Stuckey, P. J.; Guns, T.; et al. 2020 · 2020
Later among the works it cites.
Collapsing Bandits and Their Application to Public Health Intervention
Mate, A.; Killian, J. A.; Xu, H.; Perrault, A.; and Tambe, M. 2020 · 2020
Later among the works it cites.
End-to-end game-focused learning of adversary behavior in security games
Perrault, A.; Wilder, B.; Ewing, E.; Mate, A.; Dilkina, B.; and Tambe, M. 2020 · 2020
Later among the works it cites.
Poverty and shared prosperity 2020: Reversals of fortune
World Bank, . 2020 · 2020
Later among the works it cites.
A refined laser method and faster matrix multiplication
Alman, J.; and Williams, V. V. 2021 · 2021
Later among the works it cites.
Smart “predict, then optimize”
Elmachtoub, A. N.; and Grigas, P. 2021 · 2021
Later among the works it cites.
Learning MDPs from Features: Predict-Then-Optimize for Sequential Decision Making by Reinforcement Learning
Wang, K.; Shah, S.; Chen, H.; Perrault, A.; Doshi-Velez, F.; and Tambe, M. 2021 · 2021
Later among the works it cites.
ARMMAN Helping Mothers and Children
ARMMAN. 2022 · 2022
Closest in time.
Field Study in Deploying Restless Multi-Armed Bandits: Assisting Non-Profits in Improving Maternal and Child Health
Mate, A.; Madaan, L.; Taneja, A.; Madhiwalla, N.; Verma, S.; Singh, G.; Hegde, A.; Varakantham, P.; and Tambe, M. 2022 · 2022
Closest in time.