Fetching the paper…
Reading the bibliography…
Bandit algorithms have been predominantly analyzed in the convex setting with function-value based stationary regret as the performance measure.
An analysis of approximations for maximizing submodular set functions
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
An overview of the simultaneous perturbation method for efficient optimization
James C Spall · 1998
Earlier work this paper cites.
Tracking a small set of experts by mixing past posteriors
Olivier Bousquet and Manfred K Warmuth · 2002
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D Flaxman, Adam Tauman Kalai, and H Brendan McMahan · 2005
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
Rodeo: sparse, greedy nonparametric regression
John Lafferty, Larry Wasserman, et al · 2008
Earlier work this paper cites.
Efficient learning algorithms for changing environments
Elad Hazan and Comandur Seshadhri · 2009
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Alekh Agarwal and Ofer Dekel · 2010
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Gergely Neu, Andras Antos, András György, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Stochastic convex optimization with bandit feedback
Alekh Agarwal, Dean P Foster, Daniel J Hsu, Sham M Kakade, and Alexander Rakhlin · 2011
Earlier work this paper cites.
Improved regret guarantees for online smooth convex optimization with bandit feedback
Ankan Saha and Ambuj Tewari · 2011
Earlier work this paper cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
Raman Arora, Ofer Dekel, and Ambuj Tewari · 2012
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Earlier work this paper cites.
Learning with submodular functions: A convex optimization perspective
Francis Bach et al · 2013
Earlier work this paper cites.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Submodular function maximization via the multilinear relaxation and contention resolution schemes
Chandra Chekuri, Jan Vondrák, and Rico Zenklusen · 2014
Cited alongside, same era.
Online learning in markov decision processes with changing cost sequences
Travis Dick, Andras Gyorgy, and Csaba Szepesvari · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Online markov decision processes with kullback–leibler control cost
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Deep reinforcement learning: An overview
Yuxi Li · 2017
Later among the works it cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Later among the works it cites.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir · 2017
Later among the works it cites.
Krishnakumar Balasubramanian and Saeed Ghadimi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peng Guan, Maxim Raginsky, and Rebecca M Willett · 2014
Cited alongside, same era.
Non-stationary stochastic optimization
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2015
Cited alongside, same era.
Submodularity in machine learning applications
Jeff Bilmes · 2015
Cited alongside, same era.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Cited alongside, same era.
Online convex optimization in dynamic environments
Eric C Hall and Rebecca M Willett · 2015
Cited alongside, same era.
Online optimization: Competing with dynamic comparators
Ali Jadbabaie, Alexander Rakhlin, Shahin Shahrampour, and Karthik Sridharan · 2015
Cited alongside, same era.
Achieving all with no parameters: Adanormalhedge
Haipeng Luo and Robert E Schapire · 2015
Cited alongside, same era.
Lin Chen, Hamed Hassani, and Amin Karbasi · 2018
Later among the works it cites.
Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2018
Later among the works it cites.
Online learning with non-convex losses and non-stationary regret
Xiand Gao, Xiaobo Li, and Shuzhong Zhang · 2018
Later among the works it cites.
Lectures on convex optimization
Yurii Nesterov · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Stochastic cubic regularization for fast nonconvex optimization
Nilesh Tripuraneni, Mitchell Stern, Chi Jin, Jeffrey Regier, and Michael I Jordan · 2018
Later among the works it cites.
Stochastic zeroth-order optimization in high dimensions
Yining Wang, Simon Du, Sivaraman Balakrishnan, and Aarti Singh · 2018
Later among the works it cites.
Achieving optimal dynamic regret for non-stationary bandits without prior information
Peter Auer, Yifang Chen, Pratik Gajane, Chung-Wei Lee, Haipeng Luo, Ronald Ortner, and Chen-Yu Wei · 2019
Closest in time.
Submodular functions: from discrete to continuous domains
Francis Bach · 2019
Closest in time.
Online forecasting of total-variation-bounded sequences
Dheeraj Baby and Yu-Xiang Wang · 2019
Closest in time.
Black box submodular maximization: Discrete and continuous settings
Lin Chen, Mingrui Zhang, Hamed Hassani, and Amin Karbasi · 2019
Closest in time.
Effect of depth and width on local minima in deep learning
Kenji Kawaguchi, Jiaoyang Huang, and Leslie Pack Kaelbling · 2019
Closest in time.