Fetching the paper…
Reading the bibliography…
We consider contextual bandits with linear constraints (CBwLC), a variant of contextual bandits in which the algorithm consumes multiple resources subject to linear constraints on total consumption.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 1904
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 1995
Earlier work this paper cites.
Game theory, on-line prediction and boosting
Yoav Freund and Robert E Schapire · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Earlier work this paper cites.
Competitive on-line linear regression
Vladimir Vovk · 1997
Earlier work this paper cites.
Tracking the best expert
Mark Herbster and Manfred K Warmuth · 1998
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
Katy S. Azoury and Manfred K. Warmuth · 2001
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2004
Earlier work this paper cites.
Online Convex Optimization in the Bandit Setting: Gradient Descent without a Gradient
Abraham Flaxman, Adam Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Metric entropy in competitive on-line prediction
Vladimir Vovk · 2006
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms
Omar Besbes and Assaf Zeevi · 2009
Earlier work this paper cites.
Dylan J. Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual Bandits with Linear Payoff Functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Earlier work this paper cites.
Efficient optimal leanring for contextual bandits
Miroslav Dudík, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Earlier work this paper cites.
Dynamic pricing with limited supply
Moshe Babaioff, Shaddin Dughmi, Robert D. Kleinberg, and Aleksandrs Slivkins · 2012
Earlier work this paper cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Earlier work this paper cites.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Earlier work this paper cites.
Sparsity regret bounds for individual sequences in online linear regression
S. Gerchinovitz · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2014
Cited alongside, same era.
Bandits with global convex constraints and objective
Shipra Agrawal and Nikhil R. Devanur · 2014
Cited alongside, same era.
Resourceful contextual bandits
Ashwinkumar Badanidiyuru, John Langford, and Aleksandrs Slivkins · 2014
Cited alongside, same era.
Bandit convex optimization: Towards tight bounds
Elad Hazan and Kfir Y. Levy · 2014
Cited alongside, same era.
Online nonparametric regression
Alexander Rakhlin and Karthik Sridharan · 2014
Cited alongside, same era.
Adapting to misspecification in contextual bandits
Dylan J Foster, Claudio Gentile, Mehryar Mohri, and Julian Zimmert · 2020
Later among the works it cites.
Online learning with vector costs and bandits with knapsacks
Thomas Kesselheim and Sahil Singla · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
A contextual bandit bake-off
Alberto Bietti, Alekh Agarwal, and John Langford · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Later among the works it cites.
Contextual bandits in large action spaces: Made practical
Yinglun Zhu, Dylan J. Foster, Paul Mineiro, and John Langford · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Close the gaps: A learning-while-doing algorithm for single-product revenue management problems
Zizhuo Wang, Shiming Deng, and Yinyu Ye · 2014
Cited alongside, same era.
Bandit convex optimization: \(\sqrt{T}\) regret in one dimension
Sébastien Bubeck, Ofer Dekel, Tomer Koren, and Yuval Peres · 2015
Cited alongside, same era.
Linear contextual bandits with knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2016
Cited alongside, same era.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Shipra Agrawal, Nikhil R. Devanur, and Lihong Li · 2016
Cited alongside, same era.
Introduction to Online Convex Optimization
Elad Hazan · 2016
Cited alongside, same era.
A reductions approach to fair classification
Alekh Agarwal, Alina Beygelzimer, Miroslav Dudík, John Langford, and Hanna Wallach · 2017
Cited alongside, same era.
Online learning with knapsacks: the best of both worlds
Matteo Castiglioni, Andrea Celli, and Christian Kroer · 2022
Closest in time.
Provably efficient model-free constrained rl with linear function approximation
Arnob Ghosh, Xingyu Zhou, and Ness Shroff · 2022
Closest in time.
Optimal contextual bandits with knapsacks under realizibility via regression oracles
Yuxuan Han, Jialin Zeng, Yang Wang, Yang Xiang, and Jiheng Zhang · 2022
Closest in time.
Non-monotonic resource utilization in the bandits with knapsacks problem
Raunak Kumar and Robert Kleinberg · 2022
Closest in time.
Non-stationary bandits with knapsacks
Shang Liu, Jiashuo Jiang, and Xiaocheng Li · 2022
Closest in time.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2022
Closest in time.
On kernelized multi-armed bandits with constraints
Xingyu Zhou and Bo Ji · 2022
Closest in time.
Online bidding algorithms for return-on-spend constrained advertisers
Zhe Feng, Swati Padmanabhan, and Di Wang · 2023
Closest in time.
Approximately stationary bandits with knapsacks
Giannis Fikioris and Éva Tardos · 2023
Closest in time.
Budget pacing in repeated auctions: Regret and efficiency without convergence
Jason Gaitonde, Yingkai Li, Bar Light, Brendan Lucier, and Aleksandrs Slivkins · 2023
Closest in time.
No-regret algorithms in non-truthful auctions with budget and roi constraints
Gagan Aggarwal, Giannis Fikioris, and Mingfei Zhao · 2024
Closest in time.
No-regret is not enough! bandits with general constraints through adaptive regret minimization
Martino Bernasconi, Matteo Castiglioni, and Andrea Celli · 2024
Closest in time.
Online learning under budget and roi constraints via weak adaptivity
Matteo Castiglioni, Andrea Celli, and Christian Kroer · 2024
Closest in time.
Stochastic constrained contextual bandits via lyapunov optimization based estimation to decision framework
Hengquan Guo and Xin Liu · 2024
Closest in time.
Autobidders with budget and ROI constraints: Efficiency, regret, and pacing dynamics
Brendan Lucier, Sarath Pattathil, Aleksandrs Slivkins, and Mengxiao Zhang · 2024
Closest in time.