Fetching the paper…
Reading the bibliography…
We present an efficient second-order algorithm with $\tilde{O}(\frac{1}{\eta}\sqrt{T})$ regret for the bandit online multiclass problem.
The Perceptron: A probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1958
Earlier work this paper cites.
Pattern classification and scene analysis
R. O. Duda and P. E. Hart · 1973
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
J. Kivinen and M. Warmuth · 1997
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
Katy S Azoury and Manfred K Warmuth · 2001
Earlier work this paper cites.
Competitive on-line statistics
Volodya Vovk · 2001
Earlier work this paper cites.
Adaptive and self-confident on-line learning algorithms
P. Auer, N. Cesa-Bianchi, and C. Gentile · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2003
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
D. D. Lewis, Y. Yang, T. G. Rose, and F. Li · 2004
Earlier work this paper cites.
A second-order Perceptron algorithm
N. Cesa-Bianchi, A. Conconi, and C. Gentile · 2005
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Online passive-aggressive algorithms
K. Crammer, O. Dekel, J. Keshet, S. Shalev-Shwartz, and Y. Singer · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
E. Hazan, A. Agarwal, and S. Kale · 2007
Cited alongside, same era.
Efficient bandit algorithms for online multiclass prediction
S. M. Kakade, S. Shalev-Shwartz, and A. Tewari · 2008
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information
J. Langford and T. Zhang · 2008
Cited alongside, same era.
An efficient bandit algorithm for T \sqrt{T} -regret in online multiclass prediction?
J. Abernethy and A. Rakhlin · 2009
Cited alongside, same era.
Adaptive regularization of weight vectors
K. Crammer, A. Kulesza, and M. Dredze · 2009
Cited alongside, same era.
Adaptive bound optimization for online convex optimization
H Brendan McMahan and Matthew Streeter · 2010
Cited alongside, same era.
Learning to trade off between exploration and exploitation in multiclass bandit prediction
H. Valizadegan, R. Jin, and S. Wang · 2011
Later among the works it cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Later among the works it cites.
Beyond logarithmic bounds in online learning
F. Orabona, N. Cesa-Bianchi, and C. Gentile · 2012
Later among the works it cites.
Multiclass classification with bandit feedback using adaptive regularization
K. Crammer and C. Gentile · 2013
Later among the works it cites.
M. Mohri and A. Rostamizadeh · 2013
Later among the works it cites.
Taming the monster: a fast and simple algorithm for contextual bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A potential-based framework for online multi-class learning with partial feedback
S. Wang, R. Jin, and H. Valizadegan · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Efficient optimal learning for contextual bandits
M. Dudík, D. J. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang · 2011
Cited alongside, same era.
Newtron: an efficient bandit algorithm for online multiclass prediction
E. Hazan and S. Kale · 2011
Cited alongside, same era.
Online learning and online convex optimization
S. Shalev-Shwartz · 2011
Cited alongside, same era.
Efficient algorithms for adversarial contextual learning
Vasilis Syrgkanis, Akshay Krishnamurthy, and Robert Schapire
Cited in the paper.
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. E. Schapire · 2014
Later among the works it cites.
First-order regret bounds for combinatorial semi-bandits
G. Neu · 2015
Later among the works it cites.
A generalized online mirror descent with applications to classification and regression
F. Orabona, K. Crammer, and N. Cesa-Bianchi · 2015
Later among the works it cites.
BISTRO: An efficient relaxation-based method for contextual bandits
A. Rakhlin and K. Sridharan · 2016
Later among the works it cites.
Personal communication, 2017
Satyen Kale · 2017
Closest in time.