Fetching the paper…
Reading the bibliography…
We study the greedy (exploitation-only) algorithm in bandit problems with a known reward structure.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 1904
Earlier work this paper cites.
Iterated least squares in multiperiod control
T.L. Lai and Herbert Robbins · 1982
Earlier work this paper cites.
The continuum-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws incontrolled markov chains
Todd L Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2000
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W. Biermann, and Philip M. Long · 2003
Earlier work this paper cites.
Online linear optimization and adaptive routing
Baruch Awerbuch and Robert Kleinberg · 2004
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2004
Earlier work this paper cites.
Online Geometric Optimization in the Bandit Setting Against an Adaptive Adversary
H. Brendan McMahan and Avrim Blum · 2004
Earlier work this paper cites.
Online Convex Optimization in the Bandit Setting: Gradient Descent without a Gradient
Abraham Flaxman, Adam Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Improved Rates for the Stochastic Continuum-Armed Bandit Problem
Peter Auer, Ronald Ortner, and Csaba Szepesvári · 2007
Earlier work this paper cites.
The on-line shortest path problem under partial monitoring
András György, Tamás Linder, Gábor Lugosi, and György Ottucsák · 2007
Earlier work this paper cites.
Online Optimization in X-Armed Bandits
Sébastien Bubeck, Rémi Munos, Gilles Stoltz, and Csaba Szepesvari · 2008
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Combinatorial bandits
Nicolò Cesa-Bianchi and Gábor Lugosi · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Showing Relevant Ads via Lipschitz Context Multi-Armed Bandits
Tyler Lu, Dávid Pál, and Martin Pál · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N. Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Bandits, query learning, and the haystack dimension
Kareem Amin, Michael Kearns, and Umar Syed · 2011
Earlier work this paper cites.
Contextual Bandits with Linear Payoff Functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Earlier work this paper cites.
Contextual bandits with similarity information
Aleksandrs Slivkins · 2011
Earlier work this paper cites.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert E. Schapire · 2012
Cited alongside, same era.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Bayesian dynamic pricing policies: Learning and earning under a binary prior distribution
J. Michael Harrison, N. Bora Keskin, and Assaf Zeevi · 2012
Cited alongside, same era.
From bandits to experts: A tale of domination and independence
Noga Alon, Nicolò Cesa-Bianchi, Claudio Gentile, and Yishay Mansour · 2013
Cited alongside, same era.
Toward a classification of finite partial-monitoring games
András Antos, Gábor Bartók, Dávid Pál, and Csaba Szepesvári · 2013
Cited alongside, same era.
Greedy algorithm almost dominates in smoothed contextual bandits
Manish Raghavan, Aleksandrs Slivkins, Jennifer Wortman Vaughan, and Zhiwei Steven Wu · 2018
Later among the works it cites.
Reinforcement learning: Theory and algorithms, 2020
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Later among the works it cites.
Unreasonable effectiveness of greedy algorithms in multi-armed bandit with many arms
Mohsen Bayati, Nima Hamidi, Ramesh Johari, and Khashayar Khosravi · 2020
Later among the works it cites.
Structure adaptive algorithms for stochastic bandits
Rémy Degenne, Han Shao, and Wouter M. Koolen · 2020
Later among the works it cites.
Beyond UCB: optimal and efficient contextual bandits with regression oracles
Dylan J. Foster and Alexander Rakhlin · 2020
Later among the works it cites.
Crush optimism with pessimism: Structured bandits beyond asymptotic optimality
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Cited alongside, same era.
Implementing the “wisdom of the crowd”
Ilan Kremer, Yishay Mansour, and Motty Perry · 2013
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Partial monitoring - classification, regret bounds, and algorithms
Gábor Bartók, Dean P. Foster, Dávid Pál, Alexander Rakhlin, and Csaba Szepesvári · 2014
Cited alongside, same era.
Dynamic pricing with multiple products and partially specified demand distribution
Arnoud V. den Boer · 2014
Cited alongside, same era.
Simultaneously learning and optimizing using controlled variance pricing
Arnoud V. den Boer and Bert Zwart · 2014
Cited alongside, same era.
Dynamic pricing with an unknown demand model: Asymptotically optimal semi-myopic policies
N. Bora Keskin and Assaf J. Zeevi · 2014
Cited alongside, same era.
Kwang-Sung Jun and Chicheng Zhang · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
A contextual bandit bake-off
Alberto Bietti, Alekh Agarwal, and John Langford · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Optimal gradient-based algorithms for non-concave bandit optimization
Baihe Huang, Kaixuan Huang, Sham M. Kakade, Jason D. Lee, Qi Lei, Runzhe Wang, and Jiaqi Yang · 2021
Later among the works it cites.
Be greedy in multi-armed bandits, 2021
Matthieu Jedor, Jonathan Louëdec, and Vianney Perchet · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Later among the works it cites.
Bandit social learning: Exploration under myopic behavior, 2023
Kiarash Banihashem, MohammadTaghi Hajiaghayi, Suho Shin, and Aleksandrs Slivkins · 2023
Later among the works it cites.
Foundations of reinforcement learning and interactive decision making, 2023
Dylan J. Foster and Alexander Rakhlin · 2023
Later among the works it cites.
Aleksandrs Slivkins · 2023
Later among the works it cites.
Instance-optimality in interactive decision making: Toward a non-asymptotic theory
Andrew J Wagenmaker and Dylan J Foster · 2023
Later among the works it cites.
Sample complexity for quadratic bandits: Hessian dependent bounds and optimal algorithms
Qian Yu, Yining Wang, Baihe Huang, Qi Lei, and Jason D. Lee · 2023
Later among the works it cites.
Online learning in stackelberg games with an omniscient follower
Geng Zhao, Banghua Zhu, Jiantao Jiao, and Michael Jordan · 2023
Later among the works it cites.
Harnessing density ratios for online reinforcement learning
Philip Amortila, Dylan J Foster, Nan Jiang, Ayush Sekhari, and Tengyang Xie · 2024
Later among the works it cites.
Local anti-concentration class: Logarithmic regret for greedy linear contextual bandit
Seok-Jin Kim and Min-hwan Oh · 2024
Later among the works it cites.
Optimal learning for structured bandits
Bart P. G. Van Parys and Negin Golrezaei · 2024
Later among the works it cites.