Fetching the paper…
Reading the bibliography…
Bandit Convex Optimization (BCO) is a fundamental framework for modeling sequential decision-making with partial information, where the only feedback available to the player is the one-point or two-point function values.
How to use expert advice
Nicolò Cesa-Bianchi, Yoav Freund, David Haussler, David P. Helmbold, Robert E. Schapire, and Manfred K. Warmuth · 1997
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Adaptive routing with end-to-end feedback: Distributed learning and geometric approaches
Baruch Awerbuch and Robert D. Kleinberg · 2004
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert D. Kleinberg · 2004
Earlier work this paper cites.
Online geometric optimization in the bandit setting against an adaptive adversary
H. Brendan McMahan and Avrim Blum · 2004
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
The price of bandit information for online optimization
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
Efficient learning algorithms for changing environments
Elad Hazan and C. Seshadhri · 2009
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Alekh Agarwal, Ofer Dekel, and Lin Xiao · 2010
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov · 2011
Earlier work this paper cites.
Improved regret guarantees for online smooth convex optimization with bandit feedback
Ankan Saha and Ambuj Tewari · 2011
Earlier work this paper cites.
Towards minimax policies for online linear optimization with bandit feedback
Sébastien Bubeck, Nicolò Cesa-Bianchi, and Sham M. Kakade · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2012
Cited alongside, same era.
Ensemble Methods: Foundations and Algorithms
Zhi-Hua Zhou · 2012
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Yonatan Gur, Assaf J. Zeevi, and Omar Besbes · 2014
Cited alongside, same era.
Bandit convex optimization: Towards tight bounds
Elad Hazan and Kfir Y. Levy · 2014
Cited alongside, same era.
Metagrad: Multiple learning rates in online learning
Tim van Erven and Wouter M. Koolen · 2016
Later among the works it cites.
Optimistic bandit convex optimization
Scott Yang and Mehryar Mohri · 2016
Later among the works it cites.
Tracking slowly moving clairvoyant: Optimal dynamic regret of online learning with true and noisy gradient
Tianbao Yang, Lijun Zhang, Rong Jin, and Jinfeng Yi · 2016
Later among the works it cites.
Kernel-based methods for bandit convex optimization
Sébastien Bubeck, Yin Tat Lee, and Ronen Eldan · 2017
Later among the works it cites.
Improved strongly adaptive online learning using coin betting
Kwang-Sung Jun, Francesco Orabona, Stephen Wright, and Rebecca Willett · 2017
Later among the works it cites.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Non-stationary stochastic optimization
Omar Besbes, Yonatan Gur, and Assaf J. Zeevi · 2015
Cited alongside, same era.
Bandit convex optimization: T \sqrt{T} regret in one dimension
Sébastien Bubeck, Ofer Dekel, Tomer Koren, and Yuval Peres · 2015
Cited alongside, same era.
Strongly adaptive online learning
Amit Daniely, Alon Gonen, and Shai Shalev-Shwartz · 2015
Cited alongside, same era.
Bandit smooth convex optimization: Improving the bias-variance tradeoff
Ofer Dekel, Ronen Eldan, and Tomer Koren · 2015
Cited alongside, same era.
Online optimization : Competing with dynamic comparators
Ali Jadbabaie, Alexander Rakhlin, Shahin Shahrampour, and Karthik Sridharan · 2015
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2016
Cited alongside, same era.
Later among the works it cites.
Improved dynamic regret for non-degeneracy functions
Lijun Zhang, Tianbao Yang, Jinfeng Yi, Rong Jin, and Zhi-Hua Zhou · 2017
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Minimizing adaptive regret with one gradient per iteration
Guanghui Wang, Dakuan Zhao, and Lijun Zhang · 2018
Later among the works it cites.
Achieving optimal dynamic regret for non-stationary bandits without prior information
Peter Auer, Yifang Chen, Pratik Gajane, Chung-Wei Lee, Haipeng Luo, Ronald Ortner, and Chen-Yu Wei · 2019
Closest in time.
Bandit convex optimization for scalable and dynamic IoT management
Tianyi Chen and Georgios B. Giannakis · 2019
Closest in time.
Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Closest in time.
A simple approach for non-stationary linear bandits
Peng Zhao, Lijun Zhang, Yuan Jiang, and Zhi-Hua Zhou · 2020
Closest in time.