Fetching the paper…
Reading the bibliography…
This paper proposes a new method for the K-armed dueling bandit problem, a variation on the regular K-armed bandit problem that offers only relative feedback about pairs of arms.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W.R · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H · 1985
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Joachims, T · 2002
Earlier work this paper cites.
Prediction, Learning, and Games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Letor: Benchmark dataset for research on learning to rank for information retrieval
Liu, T.-Y., Xu, J., Qin, T., Xiong, W., and Li, H · 2007
Earlier work this paper cites.
An experimental comparison of click position-bias models
Craswell, N., Zoeter, O., Taylor, M., and Ramsey, B · 2008
Earlier work this paper cites.
Introduction to Information Retrieval
Manning, C., Raghavan, P., and Schütze, H · 2008
Earlier work this paper cites.
How does clickthrough data reflect retrieval quality?
Radlinski, F., Kurup, M., and Joachims, T · 2008
Earlier work this paper cites.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y., Munos, R., and Szepesvári, C · 2009
Cited alongside, same era.
Pure exploration in multi-armed bandits problems
Bubeck, S., Munos, R., and Stoltz, G · 2009
Cited alongside, same era.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yue, Y. and Joachims, T · 2009
Cited alongside, same era.
Preference Learning
Fürnkranz, J. and Hüllermeier, E. (eds.) · 2010
Cited alongside, same era.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M · 2010
Cited alongside, same era.
X-armed bandits
Bubeck, S., Munos, R., Stoltz, G., and Szepesvari, C · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Later among the works it cites.
Exponential regret bounds for Gaussian process bandits with deterministic observations
de Freitas, N., Smola, A., and Zoghi, M · 2012
Later among the works it cites.
Towards preference-based reinforcement learning
Fürnkranz, J., Hüllermeier, E., Cheng, W., and Park, S.H · 2012
Later among the works it cites.
Thompson sampling: an asymptotically optimal finite time analysis
Kauffmann, E., Korda, N., and Munos, R · 2012
Later among the works it cites.
The K-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T · 2012
Later among the works it cites.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A probabilistic method for inferring preferences from clicks
Hofmann, K., Whiteson, S., and de Rijke, M · 2011
Cited alongside, same era.
Optimistic optimization of a deterministic function without the knowledge of its smoothness
Munos, R · 2011
Cited alongside, same era.
Beat the mean bandit
Yue, Y. and Joachims, T · 2011
Cited alongside, same era.
Tailoring click models to user goals
Guo, F., Li, L., and Faloutsos, C
Cited in the paper.
Efficient multiple-click models in web search
Guo, F., Liu, C., and Wang, Y
Cited in the paper.
Cappé, O., Garivier, A., Maillard, O.-A., Munos, R., and Stoltz, G · 2013
Closest in time.
Balancing exploration and exploitation in listwise and pairwise online learning to rank for information retrieval
Hofmann, K., Whiteson, S., and de Rijke, M · 2013
Closest in time.
Generic exploration and k-armed voting bandits
Urvoy, T., Clerot, F., Féraud, R., and Naamane, S · 2013
Closest in time.
Stochastic simultaneous optimistic optimization
Valko, M., Carpentier, A., and Munos, R · 2013
Closest in time.