Fetching the paper…
Reading the bibliography…
A standard assumption in contextual multi-arm bandit is that the true context is perfectly known before arm selection.
Learning to Forget: Continual Prediction with LSTM
Gers, F. A.; Schmidhuber, J.; and Cummins, F. 2000 · 2000
Earlier work this paper cites.
Finite-time Analysis of the Multiarmed Bandit Problem
Auer, P.; Cesa-Bianchi, N.; and Fischer, P. 2002 · 2002
Earlier work this paper cites.
The Nonstochastic Multiarmed Bandit Problem
Auer, P.; Cesa-Bianchi, N.; Freund, Y.; and Schapire, R. E. 2002 · 2002
Earlier work this paper cites.
Sequential Batch Learning in Finite-Action Linear Contextual Bandits
Han, Y.; Zhou, Z.; Zhou, Z.; Blanchet, J.; Glynn, P. W.; and Ye, Y. 2020 · 2004
Earlier work this paper cites.
Wireless Communications
Goldsmith, A. 2005 · 2005
Earlier work this paper cites.
Distributional Robust Batch Contextual Bandits
Si, N.; Zhang, F.; Zhou, Z.; and Blanchet, J. 2020a · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V.; Hayes, T.; Thomas, P.; and Kakade, S. 2008 · 2008
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Audibert, J.; and Bubeck, S. 2009 · 2009
Earlier work this paper cites.
A Contextual-bandit Approach to Personalized News Article Recommendation
Li, L.; Chu, W.; Langford, J.; and Schapire, R. E. 2010 · 2010
Earlier work this paper cites.
Contextual multi-armed bandits
Lu, T.; Pál, D.; and Pál, M. 2010 · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: no regret and experimental design
Srinivas, N.; Krause, A.; Kakade, S.; and Seeger, M. 2010 · 2010
Earlier work this paper cites.
Improved Algorithms for Linear Stochastic Bandits
Abbasi-Yadkori, Y.; Pál, D.; and Szepesvári, C. 2011 · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W.; Li, L.; Reyzin, L.; and Schapire, R. 2011 · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M.; Langford, J.; and Li, L. 2011 · 2011
Earlier work this paper cites.
Dynamic right-sizing for power-proportional data centers
Lin, M.; Wierman, A.; Andrew, L. L. H.; and Thereska, E. 2011 · 2011
Earlier work this paper cites.
Analysis of Thompson Sampling for the Multi-armed Bandit Problem
Agrawal, S.; and Goyal, N. 2012 · 2012
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S.; and Cesa-Bianchi, N. 2012 · 2012
Earlier work this paper cites.
Online algorithms for geographical load balancing
Lin, M.; Liu, Z.; Wierman, A.; and Andrew, L. L. H. 2012 · 2012
Earlier work this paper cites.
Thompson Sampling for Contextual Bandits with Linear Payoffs
Agrawal, S.; and Goyal, N. 2013 · 2013
Earlier work this paper cites.
Two-stage minimax regret robust unit commitment
Jiang, R.; Wang, J.; Zhang, M.; and Guan, Y. 2013 · 2013
Cited alongside, same era.
Online Learning with Predictable Sequences
Rakhlin, A.; and Sridharan, K. 2013 · 2013
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
Valko, M.; Korda, N.; Munos, R.; Flaounas, I.; and Cristianini, N. 2013 · 2013
Cited alongside, same era.
Robust optimization for transmission expansion planning: Minimax cost vs. minimax regret
Chen, B.; Wang, J.; Wang, L.; He, Y.; and Wang, Z. 2014 · 2014
Cited alongside, same era.
An actor-critic contextual bandit algorithm for personalized interventions using mobile devices
Lei, H.; Tewari, A.; and Murphy, S. 2014 · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Garcıa, J.; and Fernández, F. 2015 · 2015
Adversarially Robust Optimization with Gaussian Processes
Bogunovic, I.; Scarlett, J.; Jegelka, S.; and Cevher, V. 2018 · 2018
Later among the works it cites.
Spatio–temporal edge service placement: A bandit learning approach
Chen, L.; Xu, J.; Ren, S.; and Zhou, P. 2018 · 2018
Later among the works it cites.
Adversarial Attacks on Stochastic Bandits
Jun, K.-S.; Li, L.; Ma, Y.; and Zhu, X. 2018 · 2018
Later among the works it cites.
Robust actor-critic contextual bandit for mobile health (mhealth) interventions
Zhu, F.; Guo, J.; Li, R.; and Huang, J. 2018 · 2018
Later among the works it cites.
Best Arm Identification for Contaminated Bandits
Altschuler, J.; Brunel, V.-E.; and Malek, A. 2019 · 2019
Later among the works it cites.
Budget-constrained edge service provisioning with demand estimation via bandit learning
Chen, L.; and Xu, J. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Online optimization: Competing with dynamic comparators
Jadbabaie, A.; Rakhlin, A.; Shahrampour, S.; and Sridharan, K. 2015 · 2015
Cited alongside, same era.
An Algorithm with Nearly Optimal Pseudo-regret for Both Stochastic and Adversarial Bandits
Auer, P.; and Chiang, C.-K. 2016 · 2016
Cited alongside, same era.
Introduction to time series and forecasting
Brockwell, P. J.; Brockwell, P. J.; Davis, R. A.; and Davis, R. A. 2016 · 2016
Cited alongside, same era.
Refined lower bounds for adversarial bandits
Gerchinovitz, S.; and Lattimore, T. 2016 · 2016
Cited alongside, same era.
Edge Computing: Vision and Challenges
Shi, W.; Cao, J.; Zhang, Q.; Li, Y.; and Xu, L. 2016 · 2016
Cited alongside, same era.
Efficient algorithms for adversarial contextual learning
Syrgkanis, V.; Krishnamurthy, A.; and Schapire, R. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Exploiting Vulnerabilities of Load Forecasting Through Adversarial Attacks
Chen, Y.; Tan, Y.; and Zhang, B. 2019 · 2019
Later among the works it cites.
Stochastic Bandits with Context Distributions
Kirschner, J.; and Krause, A. 2019 · 2019
Later among the works it cites.
Data Poisoning Attacks on Stochastic Bandits
Liu, F.; and Shroff, N. 2019 · 2019
Later among the works it cites.
Contextual Multi-Armed Bandits for Link Adaptation in Cellular Networks
Saxena, V.; Jaldén, J.; Gonzalez, J. E.; Bengtsson, M.; Tullberg, H.; and Stoica, I. 2019 · 2019
Later among the works it cites.
Introduction to Multi-Armed Bandits
Slivkins, A. 2019 · 2019
Later among the works it cites.
Robust Stochastic Bandit Algorithms under Probabilistic Unbounded Adversarial Attack
Guan, Z.; Ji, K.; Bucci Jr, D. J.; Hu, T. Y.; Palombo, J.; Liston, M.; and Liang, Y. 2020 · 2020
Later among the works it cites.
Distributionally Robust Bayesian Optimization
Kirschner, J.; Bogunovic, I.; Jegelka, S.; and Krause, A. 2020 · 2020
Later among the works it cites.
Efficient and robust algorithms for adversarial linear contextual bandits
Neu, G.; and Olkhovskaya, J. 2020 · 2020
Later among the works it cites.
Distributionally robust bayesian quadrature optimization
Nguyen, T.; Gupta, S.; Ha, H.; Rana, S.; and Venkatesh, S. 2020 · 2020
Later among the works it cites.
Robust experimentation in the continuous time bandit problem
Pourbabaee, F. 2020 · 2020
Later among the works it cites.
Amazon AWS Auto Scaling Documentation
Amazon. 2021 · 2021
Closest in time.
Robust Bandit Learning with Imperfect Context
Yang, J.; and Ren, S. 2021 · 2021
Closest in time.