Fetching the paper…
Reading the bibliography…
We propose a model for learning with bandit feedback while accounting for deterministically evolving and unobservable states that we call Bandits with Deterministically Evolving States ($B$-$DES$).
Bandit processes and dynamic allocation indices
Gittins, J. C · 1979
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Whittle, P · 1988
Earlier work this paper cites.
Online regret bounds for markov decision processes with deterministic transitions
Ortner, R · 2008
Earlier work this paper cites.
The design of approximation algorithms
Williamson, D. P. and Shmoys, D. B · 2011
Earlier work this paper cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
Dekel, O., Tewari, A., and Arora, R · 2012
Earlier work this paper cites.
Online regret bounds for undiscounted continuous reinforcement learning
Ortner, R. and Ryabko, D · 2012
Earlier work this paper cites.
Better rates for any adversarial deterministic mdp
Dekel, O. and Hazan, E · 2013
Earlier work this paper cites.
Focusing on the long-term: It’s good for users and business
Hohnhold, H., O’Brien, D., and Tang, D · 2015
Earlier work this paper cites.
Just in time recommendations: Modeling the dynamics of boredom in activity streams
Kapoor, K., Subbian, K., Srivastava, J., and Schrater, P · 2015
Cited alongside, same era.
Tight policy regret bounds for improving and decaying bandits
Heidari, H., Kearns, M. J., and Roth, A · 2016
Cited alongside, same era.
Rotting bandits
Levine, N., Crammer, K., and Mannor, S · 2017
Cited alongside, same era.
Recharging bandits
Kleinberg, R. and Immorlica, N · 2018
Cited alongside, same era.
Fighting boredom in recommender systems with linear reinforcement learning
Warlop, R., Lazaric, A., and Mary, J · 2018
Cited alongside, same era.
Blocking bandits
Basu, S., Sen, R., Sanghavi, S., and Shakkottai, S · 2019
Cited alongside, same era.
Rotting bandits are no harder than stochastic ones
Seznec, J., Locatelli, A., Carpentier, A., Lazaric, A., and Valko, M · 2019
Later among the works it cites.
Adversarial blocking bandits
Bishop, N., Chan, H., Mandal, D., and Tran-Thanh, L · 2020
Later among the works it cites.
Stochastic bandits with delay-dependent payoffs
Cella, L. and Cesa-Bianchi, N · 2020
Later among the works it cites.
Bandits with adversarial scaling
Lykouris, T., Mirrokni, V., and Leme, R. P · 2020
Later among the works it cites.
Nonstationary bandits with habituation and recovery dynamics
Mintz, Y., Aswani, A., Kaminsky, P., Flowers, E., and Fukuoka, Y · 2020
Later among the works it cites.
Contextual blocking bandits
Basu, S., Papadigenopoulos, O., Caramanis, C., and Shakkottai, S · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Esfandiari, H., Karbasi, A., Mehrabian, A., and Mirrokni, V · 2019
Cited alongside, same era.
Recovering bandits
Pike-Burke, C. and Grunewalder, S · 2019
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P
Cited in the paper.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E
Cited in the paper.
Multi-armed bandits with correlated arms
Gupta, S., Chaudhari, S., Joshi, G., and Yağan, O · 2021
Later among the works it cites.
Rebounding bandits for modeling satiation effects
Leqi, L., Kilinc Karzan, F., Lipton, Z., and Montgomery, A · 2021
Later among the works it cites.