Fetching the paper…
Reading the bibliography…
Bandits with Knapsacks (BwK) is a general model for multi-armed bandits under supply/budget constraints.
Asymptotically efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
On the-perturbation method for avoiding degeneracy
Nimrod Megiddo and R Chandrasekaran · 1988
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 1995
Earlier work this paper cites.
Introduction to linear optimization , volume 6
Dimitris Bertsimas and John N Tsitsiklis · 1997
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2000
Earlier work this paper cites.
Lectures on modern convex optimization: analysis, algorithms, and engineering applications , volume 2
Ahron Ben-Tal and Arkadi Nemirovski · 2001
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Dynamic assortment with demand learning for seasonal consumer goods
Felipe Caro and Jérémie Gallien · 2007
Earlier work this paper cites.
Continuous time associative bandit problems
András György, Levente Kocsis, Ivett Szabó, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2008
Earlier work this paper cites.
Exploration-exploitation trade-off using variance estimates in multi-armed bandits
J.-Y. Audibert, R. Munos, and Cs. Szepesvári · 2009
Earlier work this paper cites.
Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms
Omar Besbes and Assaf Zeevi · 2009
Earlier work this paper cites.
On the singularity probability of discrete random matrices
Jean Bourgain, Van H Vu, and Philip Matchett Wood · 2010
Earlier work this paper cites.
An asymptotically optimal bandit algorithm for bounded support models
Junya Honda and Akimichi Takemura · 2010
Earlier work this paper cites.
Bandits and experts in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Dynamic assortment optimization with a multinomial logit choice model and capacity constraint
Paat Rusmevichientong, Zuo-Jun Max Shen, and David B Shmoys · 2010
Earlier work this paper cites.
ϵ \epsilon -first policies for budget-limited multi-armed bandits
Long Tran-Thanh, Archie Chapman, Enrique Munoz de Cote, Alex Rogers, and Nicholas R. Jennings · 2010
Earlier work this paper cites.
Contextual Bandits with Linear Payoff Functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Cited alongside, same era.
The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond
Aurélien Garivier and Olivier Cappé · 2011
Cited alongside, same era.
A finite-time analysis of multi-armed bandits problems with kullback-leibler divergences
Odalric-Ambrym Maillard, Rémi Munos, and Gilles Stoltz · 2011
Cited alongside, same era.
Dynamic pricing with limited supply
Moshe Babaioff, Shaddin Dughmi, Robert D. Kleinberg, and Aleksandrs Slivkins · 2012
Cited alongside, same era.
Learning on a budget: posted price mechanisms for online procurement
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Yaron Singer · 2012
Cited alongside, same era.
Blind network revenue management
Omar Besbes and Assaf J. Zeevi · 2012
Cited alongside, same era.
Matroid bandits: Fast combinatorial optimization with learning
Branislav Kveton, Zheng Wen, Azin Ashkan, Hoda Eydgahi, and Brian Eriksson · 2014
Later among the works it cites.
Close the gaps: A learning-while-doing algorithm for single-product revenue management problems
Zizhuo Wang, Shiming Deng, and Yinyu Ye · 2014
Later among the works it cites.
Bandits with budgets: Regret lower bounds and optimal algorithms
Richard Combes, Chong Jiang, and Rayadurgam Srikant · 2015
Later among the works it cites.
Logarithmic regret bounds for bandits with knapsacks
Arthur Flajolet and Patrick Jaillet · 2015
Later among the works it cites.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvári · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Topics in random matrix theory , volume 132
Terence Tao · 2012
Cited alongside, same era.
Knapsack based optimal policies for budget-limited multi-armed bandits
Long Tran-Thanh, Archie Chapman, Alex Rogers, and Nicholas R. Jennings · 2012
Cited alongside, same era.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Cited alongside, same era.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Cited alongside, same era.
Combinatorial multi-armed bandit: General framework and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Cited alongside, same era.
Zheng Wen, Branislav Kveton, and Azin Ashkan · 2015
Later among the works it cites.
Algorithms with logarithmic or sublinear regret for constrained contextual bandits
Huasen Wu, R. Srikant, Xin Liu, and Chong Jiang · 2015
Later among the works it cites.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Shipra Agrawal, Nikhil R. Devanur, and Lihong Li · 2016
Later among the works it cites.
Mnl-bandit: A dynamic learning approach to assortment selection
Shipra Agrawal, Vashist Avadhanula, Vineet Goyal, and Assaf Zeevi · 2016
Later among the works it cites.
Assortment optimization under unknown multinomial logit choice models, 2017
Wang Chi Cheung and David Simchi-Levi · 2017
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Combinatorial semi-bandits with knapsacks
Karthik Abinav Sankararaman and Aleksandrs Slivkins · 2018
Later among the works it cites.
Adversarial bandits with knapsacks
Nicole Immorlica, Karthik Abinav Sankararaman, Robert Schapire, and Aleksandrs Slivkins · 2019
Later among the works it cites.
Unifying the stochastic and the adversarial bandits with knapsack
Anshuka Rangi, Massimo Franceschetti, and Long Tran-Thanh · 2019
Later among the works it cites.
Online learning with vector costs and bandits with knapsacks
Thomas Kesselheim and Sahil Singla · 2020
Closest in time.
Online allocation and pricing: Constant regret via bellman inequalities
Alberto Vera, Siddhartha Banerjee, and Itai Gurvich · 2020
Closest in time.
The symmetry between arms and knapsacks: A primal-dual approach for bandits with knapsacks
Xiaocheng Li, Chunlin Sun, and Yinyu Ye · 2021
Closest in time.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2021
Closest in time.