Fetching the paper…
Reading the bibliography…
Multi-armed bandits a simple but very powerful framework for algorithms that make decisions over time under uncertainty.
Introduction to Online Convex Optimization
Elad Hazan · 1909
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Capitalism, Socialism and Democracy
Joseph Schumpeter · 1942
Earlier work this paper cites.
Some notes on computation of games solutions
George W. Brown · 1949
Earlier work this paper cites.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
Approximation to bayes risk in repeated play
James Hannan · 1957
Earlier work this paper cites.
Behavior of sequential predictors of binary sequences
Thomas Cover · 1965
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Robert J. Aumann · 1974
Earlier work this paper cites.
Strategically zero-sum games: the class of games whose completely mixed equilibria cannot be improved upon
Herve Moulin and Jean-Paul Vial · 1978
Earlier work this paper cites.
Bandit processes and dynamic allocation indices (with discussion)
J. C. Gittins · 1979
Earlier work this paper cites.
Bandit problems: sequential allocation of experiments
Donald A. Berry and Bert Fristedt · 1985
Earlier work this paper cites.
Asymptotically efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 1991
Earlier work this paper cites.
Fractal, Chaos and Power Laws: Minutes from an Infinite Paradise
Manfred Schroeder · 1991
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K. Warmuth · 1994
Earlier work this paper cites.
The continuum-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 1995
Earlier work this paper cites.
Game theory, on-line prediction and boosting
Yoav Freund and Robert E Schapire · 1996
Earlier work this paper cites.
Market diffusion with two-sided learning
Dirk Bergemann and Juuso Välimäki · 1997
Earlier work this paper cites.
Empirical support for winnow and weighted-majority based algorithms: Results on a calendar scheduling domain
Avrim Blum · 1997
Earlier work this paper cites.
How to use expert advice
Nicolò Cesa-Bianchi, Yoav Freund, David Haussler, David P. Helmbold, Robert E. Schapire, and Manfred K. Warmuth · 1997
Earlier work this paper cites.
Calibrated learning and correlated equilibrium
Dean Foster and Rakesh Vohra · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Earlier work this paper cites.
Using and combining predictors that specialize
Yoav Freund, Robert E Schapire, Yoram Singer, and Manfred K Warmuth · 1997
Earlier work this paper cites.
Asymptotic calibration
Dean Foster and Rakesh Vohra · 1998
Earlier work this paper cites.
Concentration
Colin McDiarmid · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Regret in the on-line decision problem
Dean Foster and Rakesh Vohra · 1999
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2000
Earlier work this paper cites.
Experimentation in markets
Dirk Bergemann and Juuso Välimäki · 2000
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2000
Earlier work this paper cites.
PAC bounds for multi-armed bandit and Markov decision processes
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2002
Earlier work this paper cites.
Finding Nearest Neighbors in Growth-restricted Metrics
D.R. Karger and M. Ruhl · 2002
Earlier work this paper cites.
The Theory of Incentives: The Principal-Agent Model
Jean-Jacques Laffont and David Martimort · 2002
Earlier work this paper cites.
On sequential strategies for loss functions with memory
Neri Merhav, Erik Ordentlich, Gadiel Seroussi, and Marcelo J. Weinberger · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W. Biermann, and Philip M. Long · 2003
Earlier work this paper cites.
Online learning in online auctions
Avrim Blum, Vijay Kumar, Atri Rudra, and Felix Wu · 2003
Earlier work this paper cites.
Potential-based algorithms in on-line prediction and game theory
Nicolò Cesa-Bianchi and Gábor Lugosi · 2003
Earlier work this paper cites.
Bounded geometries, fractals, and low–distortion embeddings
Anupam Gupta, Robert Krauthgamer, and James R. Lee · 2003
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Tauman Kalai and Santosh Vempala · 2003
Earlier work this paper cites.
Price dispersion and learning in a dynamic differentiated-goods duopoly
Godfrey Keller and Sven Rady · 2003
Earlier work this paper cites.
The value of knowing a demand curve: Bounds on regret for online posted-price auctions
Robert D. Kleinberg and Frank T. Leighton · 2003
Earlier work this paper cites.
David Simchi-Levi and Yunzong Xu · 2003
Earlier work this paper cites.
Online linear optimization and adaptive routing
Baruch Awerbuch and Robert Kleinberg · 2004
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2004
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Shie Mannor and John N. Tsitsiklis · 2004
Earlier work this paper cites.
Online Geometric Optimization in the Bandit Setting Against an Adaptive Adversary
H. Brendan McMahan and Avrim Blum · 2004
Earlier work this paper cites.
Bypassing the embedding: Algorithms for low-dimensional metrics
Kunal Talwar · 2004
Earlier work this paper cites.
Name independent routing for growth bounded networks
Ittai Abraham and Dahlia Malkhi · 2005
Earlier work this paper cites.
Competition and innovation: An inverted u relationship
Philippe Aghion, Nicholas Bloom, Richard Blundell, Rachel Griffith, and Peter Howitt · 2005
Earlier work this paper cites.
Provably competitive adaptive routing
Baruch Awerbuch, David Holmer, Herbert Rubens, and Robert D. Kleinberg · 2005
Earlier work this paper cites.
From external to internal regret
Avrim Blum and Yishay Mansour · 2005
Earlier work this paper cites.
Online Convex Optimization in the Bandit Setting: Gradient Descent without a Gradient
Abraham Flaxman, Adam Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Algorithm Design
Jon Kleinberg and Eva Tardos · 2005
Earlier work this paper cites.
Triangulation and embedding using small sets of beacons
Jon Kleinberg, Aleksandrs Slivkins, and Tom Wexler · 2005
Earlier work this paper cites.
Greedy algorithm almost dominates in smoothed contextual bandits
Manish Raghavan, Aleksandrs Slivkins, Jennifer Wortman Vaughan, and Zhiwei Steven Wu · 2005
Earlier work this paper cites.
Incomplete Information and Internal Regret in Prediction of Individual Sequences
Gilles Stoltz · 2005
Earlier work this paper cites.
Hannan consistency in on-line learning in case of unbounded losses under partial monitoring
Chamy Allenberg, Peter Auer, László Györfi, and György Ottucsák · 2006
Earlier work this paper cites.
The dynamic pivot mechanism
Dirk Bergemann and Juuso Välimäki · 2006
Earlier work this paper cites.
Prediction, learning, and games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Earlier work this paper cites.
Anytime algorithms for multi-armed bandit problems
Robert Kleinberg · 2006
Earlier work this paper cites.
Bandit Based Monte-Carlo Planning
Levente Kocsis and Csaba Szepesvari · 2006
Earlier work this paper cites.
Competing bandits: The perils of exploration under competition., 2020
Guy Aridor, Yishay Mansour, Aleksandrs Slivkins, and Steven Wu · 2007
Earlier work this paper cites.
An efficient dynamic mechanism
Susan Athey and Ilya Segal · 2007
Earlier work this paper cites.
Improved Rates for the Stochastic Continuum-Armed Bandit Problem
Peter Auer, Ronald Ortner, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Dynamic assortment with demand learning for seasonal consumer goods
Felipe Caro and Jérémie Gallien · 2007
Earlier work this paper cites.
The Price of Bandit Information for Online Optimization
Varsha Dani, Thomas P. Hayes, and Sham Kakade · 2007
Earlier work this paper cites.
Continuous time associative bandit problems
András György, Levente Kocsis, Ivett Szabó, and Csaba Szepesvári · 2007
Earlier work this paper cites.
The on-line shortest path problem under partial monitoring
András György, Tamás Linder, Gábor Lugosi, and György Ottucsák · 2007
Earlier work this paper cites.
Online Learning with Prior Information
Elad Hazan and Nimrod Megiddo · 2007
Earlier work this paper cites.
CS683: Learning, Games, and Electronic Markets , a class at Cornell University
Robert Kleinberg · 2007
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Bandit algorithms for tree search
Rémi Munos and Pierre-Arnaud Coquelin · 2007
Earlier work this paper cites.
Towards fast decentralized construction of locality-aware overlay networks
Aleksandrs Slivkins · 2007
Earlier work this paper cites.
Competing in the Dark: An Efficient Algorithm for Bandit Linear Optimization
Jacob Abernethy, Elad Hazan, and Alexander Rakhlin · 2008
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
Peter L. Bartlett, Varsha Dani, Thomas Hayes, Sham Kakade, Alexander Rakhlin, and Ambuj Tewari · 2008
Earlier work this paper cites.
Regret minimization and the price of total anarchy
Avrim Blum, MohammadTaghi Hajiaghayi, Katrina Ligett, and Aaron Roth · 2008
Earlier work this paper cites.
Online Optimization in X-Armed Bandits
Sébastien Bubeck, Rémi Munos, Gilles Stoltz, and Csaba Szepesvari · 2008
Earlier work this paper cites.
Mortal multi-armed bandits
Deepayan Chakrabarti, Ravi Kumar, Filip Radlinski, and Eli Upfal · 2008
Earlier work this paper cites.
Adaptive design methods in clinical trials – a review
Shein-Chung Chow and Mark Chang · 2008
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Medium access in cognitive radio networks: A competitive multi-armed bandit framework
Lifeng Lai, Hai Jiang, and H. Vincent Poor · 2008
Earlier work this paper cites.
Dynamic pay-per-action mechanisms and applications to online advertising
Hamid Nazerzadeh, Amin Saberi, and Rakesh Vohra · 2008
Earlier work this paper cites.
Learning diverse rankings with multi-armed bandits
Filip Radlinski, Robert Kleinberg, and Thorsten Joachims · 2008
Earlier work this paper cites.
Adapting to a changing environment: the Brownian restless bandits
Aleksandrs Slivkins and Eli Upfal · 2008
Earlier work this paper cites.
An online algorithm for maximizing submodular functions
Matthew Streeter and Daniel Golovin · 2008
Earlier work this paper cites.
Beating the adaptive bandit with high probability
Jacob D. Abernethy and Alexander Rakhlin · 2009
Earlier work this paper cites.
Exploration-exploitation trade-off using variance estimates in multi-armed bandits
J.-Y. Audibert, R. Munos, and Cs. Szepesvári · 2009
Earlier work this paper cites.
Regret Bounds and Minimax Policies under Partial Monitoring
J.Y. Audibert and S. Bubeck · 2009
Earlier work this paper cites.
Characterizing truthful multi-armed bandit mechanisms
Moshe Babaioff, Yogeshwer Sharma, and Aleksandrs Slivkins · 2009
Earlier work this paper cites.
Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms
Omar Besbes and Assaf Zeevi · 2009
Earlier work this paper cites.
Combinatorial bandits
Nicolò Cesa-Bianchi and Gábor Lugosi · 2009
Earlier work this paper cites.
Regret and convergence bounds for a class of continuum-armed bandit problems
Eric W. Cope · 2009
Earlier work this paper cites.
The price of truthfulness for pay-per-click auctions
Nikhil Devanur and Sham M. Kakade · 2009
Earlier work this paper cites.
The AdWords problem: Online keyword matching with budgeted bidders under random permutations
Nikhil R. Devanur and Thomas P. Hayes · 2009
Earlier work this paper cites.
Concentration of Measure for the Analysis of Randomized Algorithms
Devdatt P. Dubhashi and Alessandro Panconesi · 2009
Earlier work this paper cites.
Online learning of assignments
Daniel Golovin, Andreas Krause, and Matthew Streeter · 2009
Earlier work this paper cites.
Better algorithms for benign bandits
Elad Hazan and Satyen Kale · 2009
Earlier work this paper cites.
Intrinsic robustness of the price of anarchy
Tim Roughgarden · 2009
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yisong Yue and Thorsten Joachims · 2009
Earlier work this paper cites.
The k-armed dueling bandits problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims · 2009
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Jean-Yves Audibert, Sébastien Bubeck, and Rémi Munos · 2010
Earlier work this paper cites.
Truthful mechanisms with implicit payment computation
Moshe Babaioff, Robert Kleinberg, and Aleksandrs Slivkins · 2010
Earlier work this paper cites.
Bandits Games and Clustering Foundations
Sébastien Bubeck · 2010
Earlier work this paper cites.
Online stochastic packing applied to display ad allocation
Jon Feldman, Monika Henzinger, Nitish Korula, Vahab S. Mirrokni, and Clifford Stein · 2010
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappé, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
An asymptotically optimal bandit algorithm for bounded support models
Junya Honda and Akimichi Takemura · 2010
Earlier work this paper cites.
Non-Stochastic Bandit Slate Problems
Satyen Kale, Lev Reyzin, and Robert E. Schapire · 2010
Earlier work this paper cites.
Sharp dichotomies for regret minimization in metric spaces
Robert Kleinberg and Aleksandrs Slivkins · 2010
Earlier work this paper cites.
Bandits and experts in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2010
Earlier work this paper cites.
Hedging structured concepts
Wouter M. Koolen, Manfred K. Warmuth, and Jyrki Kivinen · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Distributed learning in multi-armed bandit with multiple players
Keqin Liu and Qing Zhao · 2010
Earlier work this paper cites.
Showing Relevant Ads via Lipschitz Context Multi-Armed Bandits
Tyler Lu, Dávid Pál, and Martin Pál · 2010
Earlier work this paper cites.
Online Learning in Adversarial Lipschitz Environments
Odalric-Ambrym Maillard and Rémi Munos · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N. Tsitsiklis · 2010
Earlier work this paper cites.
Dynamic assortment optimization with a multinomial logit choice model and capacity constraint
Paat Rusmevichientong, Zuo-Jun Max Shen, and David B Shmoys · 2010
Earlier work this paper cites.
Ranked bandits in metric spaces: Learning optimally diverse rankings over large document collections
Aleksandrs Slivkins, Filip Radlinski, and Sreenivas Gollapudi · 2010
Cited alongside, same era.
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Cited alongside, same era.
ϵ \epsilon -first policies for budget-limited multi-armed bandits
Long Tran-Thanh, Archie Chapman, Enrique Munoz de Cote, Alex Rogers, and Nicholas R. Jennings · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Bandits, query learning, and the haystack dimension
Kareem Amin, Michael Kearns, and Umar Syed · 2011
Multiworld testing: A system for experimentation, learning, and decision-making, 2016
Alekh Agarwal, Sarah Bird, Markus Cozowicz, Miro Dudik, Luong Hoang, John Langford, Lihong Li, Dan Melamed, Gal Oshri, Siddhartha Sen, and Aleksandrs Slivkins · 2016
Later among the works it cites.
Linear contextual bandits with knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2016
Later among the works it cites.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Shipra Agrawal, Nikhil R. Devanur, and Lihong Li · 2016
Later among the works it cites.
Mnl-bandit: A dynamic learning approach to assortment selection
Shipra Agrawal, Vashist Avadhanula, Vineet Goyal, and Assaf Zeevi · 2016
Later among the works it cites.
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer and Chao-Kai Chiang · 2016
Later among the works it cites.
Economic recommendation systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributed algorithms for learning and cognitive medium access with logarithmic regret
Animashree Anandkumar, Nithin Michael, Ao Kevin Tang, and Ananthram Swami · 2011
Cited alongside, same era.
Dynamic auctions: A survey
Dirk Bergemann and Maher Said · 2011
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Cited alongside, same era.
Electrical flows, laplacian systems, and faster approximation of maximum flow in undirected graphs
Paul Christiano, Jonathan A. Kelner, Aleksander Madry, Daniel A. Spielman, and Shang-Hua Teng · 2011
Cited alongside, same era.
Contextual Bandits with Linear Payoff Functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Cited alongside, same era.
Near-optimal no-regret algorithms for zero-sum games
Constantinos Daskalakis, Alan Deckelbaum, and Anthony Kim · 2011
Cited alongside, same era.
Gal Bahar, Rann Smorodinsky, and Moshe Tennenholtz · 2016
Later among the works it cites.
Collaborative filtering with low regret
Guy Bresler, Devavrat Shah, and Luis Filipe Voloch · 2016
Later among the works it cites.
Tight (lower) bounds for the fixed budget best arm identification bandit problem
Alexandra Carpentier and Andrea Locatelli · 2016
Later among the works it cites.
Combinatorial multi-armed bandit with general reward functions
Wei Chen, Wei Hu, Fu Li, Jian Li, Yu Liu, and Pinyan Lu · 2016
Later among the works it cites.
Learning in social networks
Benjamin Golub and Evan D. Sadler · 2016
Later among the works it cites.
Towards optimal algorithms for prediction with expert advice
Nick Gravin, Yuval Peres, and Balasubramanian Sivan · 2016
Later among the works it cites.
Pricing a low-regret seller
Hoda Heidari, Mohammad Mahdian, Umar Syed, Sergei Vassilvitskii, and Sadra Yazdanbod · 2016
Later among the works it cites.
Online evaluation for information retrieval
Katja Hofmann, Lihong Li, and Filip Radlinski · 2016
Later among the works it cites.
Jointly private convex programming
Justin Hsu, Zhiyi Huang, Aaron Roth, and Zhiwei Steven Wu · 2016
Later among the works it cites.
Via: Improving internet telephony call quality using predictive relay selection
Junchen Jiang, Rajdeep Das, Ganesh Ananthanarayanan, Philip A. Chou, Venkat N. Padmanabhan, Vyas Sekar, Esbjorn Dominique, Marcin Goliszewski, Dalibor Kukoleca, Renat Vafin, and Hui Zhang · 2016
Later among the works it cites.
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Later among the works it cites.
Contextual semibandits via supervised learning oracles
Akshay Krishnamurthy, Alekh Agarwal, and Miroslav Dudík · 2016
Later among the works it cites.
Collaborative filtering bandits
Shuai Li, Alexandros Karatzoglou, and Claudio Gentile · 2016
Later among the works it cites.
Learning and efficiency in games with dynamic population
Thodoris Lykouris, Vasilis Syrgkanis, and Éva Tardos · 2016
Later among the works it cites.
Bayesian exploration: Incentivizing exploration in Bayesian games
Yishay Mansour, Aleksandrs Slivkins, Vasilis Syrgkanis, and Steven Wu · 2016
Later among the works it cites.
BISTRO: an efficient relaxation-based method for contextual bandits
Alexander Rakhlin and Karthik Sridharan · 2016
Later among the works it cites.
Multi-player bandits - a musical chairs approach
Jonathan Rosenski, Ohad Shamir, and Liran Szlak · 2016
Later among the works it cites.
Watch and learn: Optimizing from revealed preferences feedback
Aaron Roth, Jonathan Ullman, and Zhiwei Steven Wu · 2016
Later among the works it cites.
Twenty Lectures on Algorithmic Game Theory
Tim Roughgarden · 2016
Later among the works it cites.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Later among the works it cites.
A lower bound for multi-armed bandits with expert advice
Yevgeny Seldin and Gabor Lugosi · 2016
Later among the works it cites.
Best-of-k-bandits
Max Simchowitz, Kevin G. Jamieson, and Benjamin Recht · 2016
Later among the works it cites.
Online learning in repeated auctions
Jonathan Weed, Vianney Perchet, and Philippe Rigollet · 2016
Later among the works it cites.
Cascading bandits for large-scale recommendation problems
Shi Zong, Hao Ni, Kenny Sung, Nan Rosemary Ke, Zheng Wen, and Branislav Kveton · 2016
Later among the works it cites.
On frank-wolfe and equilibrium computation
Jacob D Abernethy and Jun-Kun Wang · 2017
Later among the works it cites.
Learning in repeated auctions with budgets: Regret minimization and equilibrium
Santiago R. Balseiro and Yonatan Gur · 2017
Later among the works it cites.
Mostly exploration-free algorithms for contextual bandits
Hamsa Bastani, Mohsen Bayati, and Khashayar Khosravi · 2017
Later among the works it cites.
Kernel-based methods for bandit convex optimization
Sébastien Bubeck, Yin Tat Lee, and Ronen Eldan · 2017
Later among the works it cites.
Algorithmic chaining and the role of partial feedback in online nonparametric learning
Nicolò Cesa-Bianchi, Pierre Gaillard, Claudio Gentile, and Sébastien Gerchinovitz · 2017
Later among the works it cites.
An online convex optimization approach to proactive network resource allocation
Tianyi Chen, Qing Ling, and Georgios B Giannakis · 2017
Later among the works it cites.
Assortment optimization under unknown multinomial logit choice models, 2017
Wang Chi Cheung and David Simchi-Levi · 2017
Later among the works it cites.
Minimal exploration in structured stochastic bandits
Richard Combes, Stefan Magureanu, and Alexandre Proutière · 2017
Later among the works it cites.
Oracle-efficient online learning and auction design
Miroslav Dudík, Nika Haghtalab, Haipeng Luo, Robert E. Schapire, Vasilis Syrgkanis, and Jennifer Wortman Vaughan · 2017
Later among the works it cites.
Tight lower bounds for multiplicative weights algorithmic families
Nick Gravin, Yuval Peres, and Balasubramanian Sivan · 2017
Later among the works it cites.
Learning, experimentation, and information design
Johannes Hörner and Andrzej Skrzypacz · 2017
Later among the works it cites.
Pytheas: Enabling data-driven quality of experience optimization using group-based exploration-exploitation
Junchen Jiang, Shijie Sun, Vyas Sekar, and Hui Zhang · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
A survey of algorithms and analysis for adaptive online learning
H. Brendan McMahan · 2017
Later among the works it cites.
Online convex optimization with time-varying constraints
Michael J Neely and Hao Yu · 2017
Later among the works it cites.
An experimental evaluation of regret-based econometrics
Noam Nisan and Gali Noti · 2017
Later among the works it cites.
Multidimensional dynamic pricing for welfare maximization
Aaron Roth, Aleksandrs Slivkins, Jonathan Ullman, and Zhiwei Steven Wu · 2017
Later among the works it cites.
An improved parametrization and analysis of the EXP3++ algorithm for stochastic and adversarial bandits
Yevgeny Seldin and Gábor Lugosi · 2017
Later among the works it cites.
Safety-aware algorithms for adversarial contextual bandit
Wen Sun, Debadeepta Dey, and Ashish Kapoor · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miroslav Dudík, John Langford, Damien Jose, and Imed Zitouni · 2017
Later among the works it cites.
Multiplicative weights update in zero-sum games
James P. Bailey and Georgios Piliouras · 2018
Later among the works it cites.
Crowdsourcing exploration
Kostas Bimpikis, Yiangos Papanastasiou, and Nicos Savva · 2018
Later among the works it cites.
Selling to a no-regret buyer
Mark Braverman, Jieming Mao, Jon Schneider, and Matt Weinberg · 2018
Later among the works it cites.
Online saddle point problem with applications to constrained online convex optimization
Adrian Rivera Cardoso, He Wang, and Huan Xu · 2018
Later among the works it cites.
Incentivizing exploration by heterogeneous users
Bangrui Chen, Peter I. Frazier, and David Kempe · 2018
Later among the works it cites.
Bandit convex optimization for scalable and dynamic iot management
Tianyi Chen and Georgios B Giannakis · 2018
Later among the works it cites.
Training gans with optimism
Constantinos Daskalakis, Andrew Ilyas, Vasilis Syrgkanis, and Haoyang Zeng · 2018
Later among the works it cites.
PCC vivace: Online-learning congestion control
Mo Dong, Tong Meng, Doron Zarchy, Engin Arslan, Yossi Gilad, Brighten Godfrey, and Michael Schapira · 2018
Later among the works it cites.
Learning to bid without knowing your value
Zhe Feng, Chara Podimata, and Vasilis Syrgkanis · 2018
Later among the works it cites.
Practical contextual bandits with regression oracles
Dylan J. Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert E. Schapire · 2018
Later among the works it cites.
Instrument-armed bandits
Nathan Kallus · 2018
Later among the works it cites.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
Sampath Kannan, Jamie Morgenstern, Aaron Roth, Bo Waggoner, and Zhiwei Steven Wu · 2018
Later among the works it cites.
Preventing fairness gerrymandering: Auditing and learning for subgroup fairness
Michael Kearns, Seth Neel, Aaron Roth, and Zhiwei Steven Wu · 2018
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Competing bandits: Learning under competition
Yishay Mansour, Aleksandrs Slivkins, and Steven Wu · 2018
Later among the works it cites.
Cycles in adversarial regularized learning
Panayotis Mertikopoulos, Christos H. Papadimitriou, and Georgios Piliouras · 2018
Later among the works it cites.
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
Combinatorial semi-bandits with knapsacks
Karthik Abinav Sankararaman and Aleksandrs Slivkins · 2018
Later among the works it cites.
Human interaction with recommendation systems
Sven Schmit and Carlos Riquelme · 2018
Later among the works it cites.
Acceleration through optimistic no-regret dynamics
Jun-Kun Wang and Jacob D. Abernethy · 2018
Later among the works it cites.
More adaptive algorithms for adversarial bandits
Chen-Yu Wei and Haipeng Luo · 2018
Later among the works it cites.
Reinforcement learning: Theory and algorithms, 2020
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Closest in time.
The perils of exploration under competition: A computational modeling approach
Guy Aridor, Kevin Liu, Aleksandrs Slivkins, and Steven Wu · 2019
Closest in time.
Adaptively tracking the best arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2019
Closest in time.
Social learning and the innkeeper’s challenge
Gal Bahar, Rann Smorodinsky, and Moshe Tennenholtz · 2019
Closest in time.
Information design: A unified perspective
Dirk Bergemann and Stephen Morris · 2019
Closest in time.
SIC-MMAB: synchronisation involves communication in multiplayer multi-armed bandits
Etienne Boursier and Vianney Perchet · 2019
Closest in time.
Multi-armed bandit problems with strategic arms
Mark Braverman, Jieming Mao, Jon Schneider, and S. Matthew Weinberg · 2019
Closest in time.
Improved path-length regret bounds for bandits
Sébastien Bubeck, Yuanzhi Li, Haipeng Luo, and Chen-Yu Wei · 2019
Closest in time.
A new algorithm for non-stationary contextual bandits: Efficient, optimal, and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Closest in time.
Vortices instead of equilibria in minmax optimization: Chaos and butterfly effects of online learning in zero-sum games
Yun Kuen Cheung and Georgios Piliouras · 2019
Closest in time.
Optimal algorithm for bayesian incentive-compatible exploration
Lee Cohen and Yishay Mansour · 2019
Closest in time.
Last-iterate convergence: Zero-sum games and constrained min-max optimization
Constantinos Daskalakis and Ioannis Panageas · 2019
Closest in time.
Better algorithms for stochastic bandits with adversarial corruptions
Anupam Gupta, Tomer Koren, and Kunal Talwar · 2019
Closest in time.
Bayesian exploration with heterogenous agents
Nicole Immorlica, Jieming Mao, Aleksandrs Slivkins, and Steven Wu · 2019
Closest in time.
Adversarial bandits with knapsacks
Nicole Immorlica, Karthik Abinav Sankararaman, Robert Schapire, and Aleksandrs Slivkins · 2019
Closest in time.
Bayesian persuasion and information design
Emir Kamenica · 2019
Closest in time.
Contextual bandits with continuous actions: Smoothing, zooming, and adapting
Akshay Krishnamurthy, John Langford, Aleksandrs Slivkins, and Chicheng Zhang · 2019
Closest in time.
Unifying the stochastic and the adversarial bandits with knapsack
Anshuka Rangi, Massimo Franceschetti, and Long Tran-Thanh · 2019
Closest in time.
Personal communication
Mark Sellke, 2019 · 2019
Closest in time.
Connections between mirror descent, thompson sampling and the information ratio
Julian Zimmert and Tor Lattimore · 2019
Closest in time.
An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2019
Closest in time.
Beating stochastic and adversarial semi-bandits optimally and simultaneously
Julian Zimmert, Haipeng Luo, and Chen-Yu Wei · 2019
Closest in time.
Fiduciary bandits
Gal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, and Moshe Tennenholtz · 2020
Closest in time.
Unreasonable effectiveness of greedy algorithms in multi-armed bandit with many arms
Mohsen Bayati, Nima Hamidi, Ramesh Johari, and Khashayar Khosravi · 2020
Closest in time.
First-order bayesian regret analysis of thompson sampling
Sébastien Bubeck and Mark Sellke · 2020
Closest in time.
Non-stochastic multi-player multi-armed bandits: Optimal rate with collision information, sublinear without
Sébastien Bubeck, Yuanzhi Li, Yuval Peres, and Mark Sellke · 2020
Closest in time.
Beyond UCB: optimal and efficient contextual bandits with regression oracles
Dylan J. Foster and Alexander Rakhlin · 2020
Closest in time.
Tight last-iterate convergence rates for no-regret learning in multi-player games
Noah Golowich, Sarath Pattathil, and Constantinos Daskalakis · 2020
Closest in time.
Incentivizing exploration with selective data disclosure
Nicole Immorlica, Jieming Mao, Aleksandrs Slivkins, and Steven Wu · 2020
Closest in time.
Online learning with vector costs and bandits with knapsacks
Thomas Kesselheim and Sahil Singla · 2020
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
Efficient contextual bandits with continuous actions
Maryam Majzoubi, Chicheng Zhang, Rajan Chari, Akshay Krishnamurthy, John Langford, and Aleksandrs Slivkins · 2020
Closest in time.
Bandits with knapsacks beyond the worst-case, 2021
Karthik Abinav Sankararaman and Aleksandrs Slivkins · 2020
Closest in time.
Online allocation and pricing: Constant regret via bellman inequalities
Alberto Vera, Siddhartha Banerjee, and Itai Gurvich · 2020
Closest in time.
Mnl-bandit with knapsacks
Abdellah Aznag, Vineet Goyal, and Noemie Perivier · 2021
Closest in time.
A contextual bandit bake-off
Alberto Bietti, Alekh Agarwal, and John Langford · 2021
Closest in time.
Be greedy in multi-armed bandits, 2021
Matthieu Jedor, Jonathan Louëdec, and Vianney Perchet · 2021
Closest in time.
The symmetry between arms and knapsacks: A primal-dual approach for bandits with knapsacks
Xiaocheng Li, Chunlin Sun, and Yinyu Ye · 2021
Closest in time.
Incentivizing compliance with algorithmic instrumentsincentivizing compliance with algorithmic instruments
Daniel Ngo, Logan Stapleton, Vasilis Syrgkanis, and Steven Wu · 2021
Closest in time.
Adaptive discretization for adversarial lipschitz bandits
Chara Podimata and Aleksandrs Slivkins · 2021
Closest in time.
The price of incentivizing exploration: A characterization via thompson sampling and sample complexity
Mark Sellke and Aleksandrs Slivkins · 2021
Closest in time.
Incentives and exploration in reinforcement learning
Max Simchowitz and Aleksandrs Slivkins · 2021
Closest in time.
Linear last-iterate convergence in constrained saddle-point optimization
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, and Haipeng Luo · 2021
Closest in time.
Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2021
Closest in time.
Optimal contextual bandits with knapsacks under realizibility via regression oracles
Yuxuan Han, Jialin Zeng, Yang Wang, Yang Xiang, and Jiheng Zhang · 2022
Closest in time.
Bandit social learning: Exploration under myopic behavior
Kiarash Banihashem, MohammadTaghi Hajiaghayi, Suho Shin, and Aleksandrs Slivkins · 2023
Closest in time.
Exploration and persuasion
Aleksandrs Slivkins · 2023
Closest in time.
Contextual bandits with packing and covering constraints: A modular lagrangian approach via regression
Aleksandrs Slivkins, Karthik Abinav Sankararaman, and Dylan J. Foster · 2023
Closest in time.