Fetching the paper…
Reading the bibliography…
We study structured multi-armed bandits, which is the problem of online decision-making under uncertainty in the presence of structural information.
Thompson W (1933) On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25(3/4):285–294
1933
Earlier work this paper cites.
Bank B, Guddat J, Klatte D, Kummer B, Tammer K (1982) Non-Linear Parametric Optimization (Springer)
1982
Earlier work this paper cites.
Lai T, Robbins H (1985) Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6(1):4–22
1985
Earlier work this paper cites.
Agrawal R, Teneketzis D, Anantharam V (1989) Asymptotically efficient adaptive allocation schemes for controlled iid processes: Finite parameter space. IEEE Transactions on Automatic Control 34(3)
1989
Earlier work this paper cites.
Hoeffding W (1994) Probability inequalities for sums of bounded random variables. The Collected Works of Wassily Hoeffding , 409–426 (Springer)
1994
Earlier work this paper cites.
Berge C (1997) Topological Spaces: Including a Treatment of Multi-Valued Functions, Vector Spaces, and Convexity (Courier Corporation)
1997
Earlier work this paper cites.
Graves T, Lai T (1997) Asymptotically efficient adaptive choice of control laws incontrolled Markov chains. SIAM Journal on Control and Optimization 35(3):715–743
1997
Earlier work this paper cites.
Barvinok A (2002) A Course in Convexity , volume 54 (American Mathematical Society)
2002
Earlier work this paper cites.
Boyd S, Vandenberghe L (2004) Convex Optimization (Cambridge University Press)
2004
Earlier work this paper cites.
Nesterov Y (2004) Introductory Lectures on Convex Optimization: A Basic Course (Kluwer Academic Publishers)
2004
Earlier work this paper cites.
Dani V, Hayes T, Kakade S (2008) Stochastic linear optimization under bandit feedback. International Conference on Algorithmic Learning Theory , 355–366
2008
Earlier work this paper cites.
Streeter M, Golovin D (2008) An online algorithm for maximizing submodular functions. Advances in Neural Information Processing Systems , 1577–1584
2008
Earlier work this paper cites.
Aubin JP, Frankowska H (2009) Set-Valued Analysis (Springer Science & Business Media)
2009
Earlier work this paper cites.
Bertsekas D (2009) Convex Optimization Theory (Athena Scientific Belmont)
2009
Earlier work this paper cites.
Besbes O, Zeevi A (2009) Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms. Operations Research 57(6):1407–1420
2009
Cited alongside, same era.
Mersereau A, Rusmevichientong P, Tsitsiklis J (2009) A structured multiarmed bandit problem and the greedy policy. IEEE Transactions on Automatic Control 54(12):2787–2802
2009
Cited alongside, same era.
Filippi S, Cappe O, Garivier A, Szepesvári C (2010) Parametric bandits: The generalized linear case. NIPS , volume 23, 586–594
2010
Cited alongside, same era.
Rusmevichientong P, Tsitsiklis JN (2010) Linearly parameterized bandits. Mathematics of Operations Research 35(2):395–411
2010
Cited alongside, same era.
Garivier A, Cappé O (2011) The KL-UCB algorithm for bounded stochastic bandits and beyond. Proceedings of the 24th annual conference on learning theory , 359–376
Chandrasekaran V, Shah P (2017) Relative entropy optimization and its applications. Mathematical Programming 161(1-2):1–32
2017
Later among the works it cites.
Combes R, Magureanu S, Proutiere A (2017) Minimal exploration in structured stochastic bandits. Advances in Neural Information Processing Systems , 1763–1771
2017
Later among the works it cites.
Lattimore T, Szepesvari C (2017) The end of optimism? an asymptotic analysis of finite-armed linear bandits. volume 54 of Proceedings of Machine Learning Research , 728–737 (Fort Lauderdale, FL, USA: PMLR)
2017
Later among the works it cites.
Combettes P (2018) Perspective functions: Properties, constructions, and examples. Set-Valued and Variational Analysis 26(2):247–264
2018
Later among the works it cites.
Mao J, Leme R, Schneider J (2018) Contextual pricing for Lipschitz buyers. Advances in Neural Information Processing Systems , 5643–5651
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
Slivkins A (2011) Contextual bandits with similarity information. Proceedings of the 24th annual Conference On Learning Theory , 679–702
2011
Cited alongside, same era.
Bubeck S, Cesa-Bianchi N (2012) Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations and Trends® in Machine Learning 5(1):1–122
2012
Cited alongside, same era.
Cover T, Thomas J (2012) Elements of Information Theory (John Wiley & Sons)
2012
Cited alongside, same era.
Kaufmann E, Korda N, Munos R (2012) Thompson sampling: An asymptotically optimal finite-time analysis. International Conference on Algorithmic Learning Theory , 199–213 (Springer)
2012
Cited alongside, same era.
Cappé O, Garivier A, Maillard OA, Munos R, Stoltz G (2013) Kullback-leibler upper confidence bounds for optimal sequential allocation. The Annals of Statistics 41(3):1516–1541
2013
Cited alongside, same era.
Combes R, Proutiere A (2014) Unimodal bandits: Regret lower bounds and optimal algorithms. International Conference on Machine Learning , 521–529
2014
Cited alongside, same era.
Keskin N, Zeevi A (2014) Dynamic pricing with an unknown demand model: Asymptotically optimal semi-myopic policies. Operations Research 62(5):1142–1167
2014
Cited alongside, same era.
2018
Later among the works it cites.
Russo D, Van Roy B (2018) Learning to optimize via information-directed sampling. Operations Research 66(1):230–252
2018
Later among the works it cites.
Balseiro S, Golrezaei N, Mahdian M, Mirrokni V, Schneider J (2019) Contextual bandits with cross-learning. Advances in Neural Information Processing Systems , 9676–9685
2019
Later among the works it cites.
Bubeck S, Devanur N, Huang Z, Niazadeh R (2019) Multi-scale online learning: Theory and applications to online auctions and pricing. Journal of Machine Learning Research
2019
Later among the works it cites.
Golrezaei N, Javanmard A, Mirrokni V (2019) Dynamic incentive-aware learning: Robust pricing in contextual auctions. Advances in Neural Information Processing Systems , 9756–9766
2019
Later among the works it cites.
Gupta S, Chaudhari S, Mukherjee S, Joshi G, Yağan O (2020) A unified approach to translate classical bandit algorithms to the structured bandit setting. IEEE Journal on Selected Areas in Information Theory 1(3):840–853
2020
Closest in time.
Jun KS, Zhang C (2020) Crush optimism with pessimism: Structured bandits beyond asymptotic optimality. Advances in Neural Information Processing Systems 33
2020
Closest in time.
Niazadeh R, Golrezaei N, Wang J, Susan F, Badanidiyuru A (2020) Online learning via offline greedy: Applications in market design and optimization. Available at SSRN 3613756
2020
Closest in time.
2020
Closest in time.
Gupta S, Chaudhari S, Joshi G, Yağan O (2021) Multi-armed bandits with correlated arms. IEEE Transactions on Information Theory 67(10):6711–6732
2021
Closest in time.