Fetching the paper…
Reading the bibliography…
In this paper, we propose a Double Thompson Sampling (D-TS) algorithm for dueling bandit problems.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Y. Yue and T. Joachims · 2009
Earlier work this paper cites.
Bandits games and clustering foundations
S. Bubeck · 2010
Earlier work this paper cites.
http://research.microsoft.com/en-us/projects/mslr/, 2010
Microsoft Research, Microsoft Learning to Rank Datasets · 2010
Earlier work this paper cites.
An empirical evaluation of Thompson Sampling
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
Beat the mean bandit
Y. Yue and T. Joachims · 2011
Earlier work this paper cites.
The k k -armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2012
Earlier work this paper cites.
Analysis of Thompson Sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Earlier work this paper cites.
Thompson sampling for the dueling bandits problem
N. Welsh · 2012
Cited alongside, same era.
Further optimal regret bounds for Thompson Sampling
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Generic exploration and k-armed voting bandits
T. Urvoy, F. Clerot, R. Féraud, and S. Naamane · 2013
Cited alongside, same era.
Relative confidence sampling for efficient on-line ranker evaluation
M. Zoghi, S. A. Whiteson, M. De Rijke, and R. Munos · 2014
Cited alongside, same era.
Relative upper confidence bound for the k k -armed dueling bandit problem
M. Zoghi, S. Whiteson, R. Munos, and M. D. Rijke · 2014
Cited alongside, same era.
Thompson sampling for complex online problems
A. Gopalan, S. Mannor, and Y. Mansour · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Later among the works it cites.
Regret lower bound and optimal algorithm in dueling bandit problem
J. Komiyama, J. Honda, H. Kashima, and H. Nakagawa · 2015
Later among the works it cites.
Copeland dueling bandits
M. Zoghi, Z. S. Karnin, S. Whiteson, and M. de Rijke · 2015
Later among the works it cites.
Optimal regret analysis of Thompson Sampling in stochastic multi-armed bandit problem with multiple plays
J. Komiyama, J. Honda, and H. Nakagawa · 2015
Later among the works it cites.
Thompson sampling for budgeted multi-armed bandits
Y. Xia, H. Li, T. Qin, N. Yu, and T.-Y. Liu · 2015
Later among the works it cites.
Thompson sampling for learning parameterized Markov decision processes
A. Gopalan and S. Mannor · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reducing dueling bandits to cardinal bandits
N. Ailon, Z. Karnin, and T. Joachims · 2014
Cited alongside, same era.
An information-theoretic analysis of Thompson Sampling
D. Russo and B. Van Roy · 2014
Cited alongside, same era.
Sparse dueling bandits
K. Jamieson, S. Katariya, A. Deshpande, and R. Nowak · 2015
Later among the works it cites.
Copeland dueling bandit problem: Regret lower bound, optimal algorithm, and computationally efficient algorithm
J. Komiyama, J. Honda, and H. Nakagawa · 2016
Closest in time.