Fetching the paper…
Reading the bibliography…
Multi-armed bandit problems are the predominant theoretical model of exploration-exploitation tradeoffs in learning, and they have countless applications ranging from medical trials, to communication networks, to Web search and advertising.
Multi-armed bandits and the Gittins index
Peter Whittle · 1980
Earlier work this paper cites.
Asymptotically efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K. Warmuth · 1994
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 1995
Earlier work this paper cites.
Fast approximation algorithms for fractional packing and covering problems
Serge A. Plotkin, David B. Shmoys, and Eva Tardos · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Earlier work this paper cites.
The complexity of optimal queuing network control
Christos H. Papadimitriou and John N. Tsitsiklis · 1999
Earlier work this paper cites.
PAC bounds for multi-armed bandit and Markov decision processes
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2002
Earlier work this paper cites.
Online learning in online auctions
Avrim Blum, Vijay Kumar, Atri Rudra, and Felix Wu · 2003
Earlier work this paper cites.
The value of knowing a demand curve: Bounds on regret for online posted-price auctions
Robert Kleinberg and Tom Leighton · 2003
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2004
Earlier work this paper cites.
Faster and simpler algorithms for multicommodity flow and other fractional packing problems
Naveen Garg and Jochen Könemann · 2007
Earlier work this paper cites.
Multi-armed Bandits with Metric Switching Costs
Sudipta Guha and Kamesh Munagala · 2007
Earlier work this paper cites.
Continuous time associative bandit problems
András György, Levente Kocsis, Ivett Szabó, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Online Learning with Prior Information
Elad Hazan and Nimrod Megiddo · 2007
Earlier work this paper cites.
Lecture notes for CS 683 (week 2), Cornell University, 2007
Robert Kleinberg · 2007
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2008
Cited alongside, same era.
Characterizing truthful multi-armed bandit mechanisms
Moshe Babaioff, Yogeshwer Sharma, and Aleksandrs Slivkins · 2009
Cited alongside, same era.
Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms
Omar Besbes and Assaf Zeevi · 2009
Cited alongside, same era.
The price of truthfulness for pay-per-click auctions
Nikhil Devanur and Sham M. Kakade · 2009
Cited alongside, same era.
The AdWords problem: Online keyword matching with budgeted bidders under random permutations
Nikhil R. Devanur and Thomas P. Hayes · 2009
Cited alongside, same era.
Online stochastic packing applied to display ad allocation
Jon Feldman, Monika Henzinger, Nitish Korula, Vahab S. Mirrokni, and Clifford Stein · 2010
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Later among the works it cites.
Geometry of online packing linear programs
Marco Molinaro and R. Ravi · 2012
Later among the works it cites.
Knapsack based optimal policies for budget-limited multi-armed bandits
Long Tran-Thanh, Archie Chapman, Alex Rogers, and Nicholas R. Jennings · 2012
Later among the works it cites.
Adaptive crowdsourcing algorithms for the bandit survey problem
Ittai Abraham, Omar Alonso, Vasilis Kandylas, and Aleksandrs Slivkins · 2013
Closest in time.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Closest in time.
Regret minimization for reserve prices in second-price auctions
Nicoló Cesa-Bianchi, Claudio Gentile, and Yishay Mansour · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sharp dichotomies for regret minimization in metric spaces
Robert Kleinberg and Aleksandrs Slivkins · 2010
Cited alongside, same era.
Showing Relevant Ads via Lipschitz Context Multi-Armed Bandits
Tyler Lu, Dávid Pál, and Martin Pál · 2010
Cited alongside, same era.
ϵ \epsilon -first policies for budget-limited multi-armed bandits
Long Tran-Thanh, Archie Chapman, Enrique Munoz de Cote, Alex Rogers, and Nicholas R. Jennings · 2010
Cited alongside, same era.
Contextual Bandits with Linear Payoff Functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Cited alongside, same era.
Near optimal online algorithms and fast approximation algorithms for resource allocation problems
Nikhil R. Devanur, Kamal Jain, Balasubramanian Sivan, and Christopher A. Wilkens · 2011
Cited alongside, same era.
Efficient optimal leanring for contextual bandits
Miroslav Dudíik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Multi-armed bandit with budget constraint and variable costs
Wenkui Ding, Tao Qin, Xu-Dong Zhang, and Tie-Yan Liu · 2013
Closest in time.
Truthful incentives in crowdsourcing tasks using regret minimization mechanisms
Adish Singla and Andreas Krause · 2013
Closest in time.
Online decision making in crowdsourcing markets: Theoretical challenges
Aleksandrs Slivkins and Jennifer Wortman Vaughan · 2013
Closest in time.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Closest in time.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2014
Closest in time.
A dynamic near-optimal algorithm for online linear programming
Shipra Agrawal, Zizhuo Wang, and Yinyu Ye · 2014
Closest in time.
Resourceful contextual bandits
Ashwinkumar Badanidiyuru, John Langford, and Aleksandrs Slivkins · 2014
Closest in time.
Efficient regret bounds for online bid optimisation in budget-limited sponsored search auctions
Long Tran-Thanh, Lampros C. Stavrogiannis, Victor Naroditskiy, Valentin Robu, Nicholas R. Jennings, and Peter Key · 2014
Closest in time.
Close the gaps: A learning-while-doing algorithm for single-product revenue management problems
Zizhuo Wang, Shiming Deng, and Yinyu Ye · 2014
Closest in time.
Dynamic pricing and learning: Historical origins, current research, and new directions
Arnoud V. Den Boer · 2015
Closest in time.
Linear contextual bandits with knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2016
Closest in time.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Shipra Agrawal, Nikhil R. Devanur, and Lihong Li · 2016
Closest in time.