Fetching the paper…
Reading the bibliography…
We study the stochastic contextual bandit with knapsacks (CBwK) problem, where each action, taken upon a context, not only leads to a random reward but also costs a random resource consumption in a vector form.
Associative reinforcement learning using linear probabilistic concepts
Abe, N. and Long, P. M. (1999) · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002) · 2002
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
Langford, J. and Zhang, T. (2007) · 2007
Earlier work this paper cites.
Smooth contextual bandits: Bridging the parametric and non-differentiable regret regimes
Hu, Y., Kallus, N., and Mao, X. (2020) · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
Rigollet, P. and Zeevi, A. (2010) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. (2011) · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Dudik, M., Hsu, D., Kale, S., Karampatziakis, N., Langford, J., Reyzin, L., and Zhang, T. (2011) · 2011
Earlier work this paper cites.
Contextual bandits with similarity information
Slivkins, A. (2011) · 2011
Earlier work this paper cites.
Trading regret for efficiency: online convex optimization with long term constraints
Mahdavi, M., Jin, R., and Yang, T. (2012) · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K. (2012) · 2012
Earlier work this paper cites.
Online learning and online convex optimization
Shalev-Shwartz, S. et al. (2012) · 2012
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., and Schapire, R. (2014) · 2014
Earlier work this paper cites.
Resourceful contextual bandits
Badanidiyuru, A., Langford, J., and Slivkins, A. (2014) · 2014
Cited alongside, same era.
Linear contextual bandits with knapsacks
Agrawal, S. and Devanur, N. (2016) · 2016
Cited alongside, same era.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Agrawal, S., Devanur, N. R., and Li, L. (2016) · 2016
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E. et al. (2016) · 2016
Cited alongside, same era.
Adaptive algorithms for online convex optimization with long-term constraints
Jenatton, R., Huang, J., and Archambeau, C. (2016) · 2016
Cited alongside, same era.
Contextual semibandits via supervised learning oracles
Krishnamurthy, A., Agarwal, A., and Dudik, M. (2016) · 2016
Cited alongside, same era.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Later among the works it cites.
A contextual bandit bake-off
Bietti, A., Agarwal, A., and Langford, J. (2021) · 2021
Later among the works it cites.
Network revenue management with nonparametric demand learning: \ \backslash sqrt { \{ T } \} -regret and polynomial dimension dependency
Miao, S. and Wang, Y. (2021) · 2021
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
Simchi-Levi, D. and Xu, Y. (2021) · 2021
Later among the works it cites.
On the re-solving heuristic for (binary) contextual bandits with knapsacks
Ai, R., Chen, Z., Deng, X., Pan, Y., Wang, C., and Yang, M. (2022) · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Empirical entropy, minimax regret and minimax risk
Rakhlin, A., Sridharan, K., and Tsybakov, A. B. (2017) · 2017
Cited alongside, same era.
Bandits with knapsacks
Badanidiyuru, A., Kleinberg, R., and Slivkins, A. (2018) · 2018
Cited alongside, same era.
Online network revenue management using thompson sampling
Ferreira, K. J., Simchi-Levi, D., and Wang, H. (2018) · 2018
Cited alongside, same era.
Practical contextual bandits with regression oracles
Foster, D., Agarwal, A., Dudik, M., Luo, H., and Schapire, R. (2018) · 2018
Cited alongside, same era.
Introduction to multi-armed bandits
Slivkins, A. et al. (2019) · 2019
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Foster, D. and Rakhlin, A. (2020) · 2020
Cited alongside, same era.
Balseiro, S. R., Lu, H., and Mirrokni, V. (2022) · 2022
Closest in time.
Online learning with knapsacks: the best of both worlds
Castiglioni, M., Celli, A., and Kroer, C. (2022) · 2022
Closest in time.
Fairness-aware network revenue management with demand learning
Chen, X., Lyu, J., Wang, Y., and Zhou, Y. (2022) · 2022
Closest in time.
Smart “predict, then optimize”
Elmachtoub, A. N. and Grigas, P. (2022) · 2022
Closest in time.
Online contextual decision-making with a smart predict-then-optimize method
Liu, H. and Grigas, P. (2022) · 2022
Closest in time.
Smoothed adversarial linear contextual bandits with knapsacks
Sivakumar, V., Zuo, S., and Banerjee, A. (2022) · 2022
Closest in time.
Efficient contextual bandits with knapsacks via regression
Slivkins, A. and Foster, D. (2022) · 2022
Closest in time.
Contextual bandits with large action spaces: Made practical
Zhu, Y., Foster, D. J., Langford, J., and Mineiro, P. (2022) · 2022
Closest in time.