Fetching the paper…
Reading the bibliography…
We present algorithms for reducing the Dueling Bandits problem to the conventional (stochastic) Multi-Armed Bandits problem.
Some aspects of the sequential design of experiments
Robbins, H · 1952
Earlier work this paper cites.
Computing with noisy information
Feige, Uriel, Raghavan, Prabhakar, Peleg, David, and Upfal, Eli · 1994
Earlier work this paper cites.
Large margin rank boundaries for ordinal regression
Herbrich, R, Graepel, Thore, and Obermayer, Klaus · 2000
Earlier work this paper cites.
Discrete prediction games with arbitrary feedback and loss
Piccolboni, Antonio and Schindelhauer, Christian · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, Peter, Cesa-Bianchi, Nicolò, and Fischer, Paul · 2002
Earlier work this paper cites.
An efficient boosting algorithm for combining preferences
Freund, Yoav, Iyer, Raj D., Schapire, Robert E., and Singer, Yoram · 2003
Earlier work this paper cites.
Evaluating the accuracy of implicit feedback from clicks and query reformulations in web search
Joachims, T., Granka, L., Pan, Bing, Hembrooke, H., Radlinski, F., and Gay, G · 2007
Earlier work this paper cites.
Noisy binary search and its applications
Karp, Richard M. and Kleinberg, Robert · 2007
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Dani, Varsha, Hayes, Thomas P., and Kakade, Sham M · 2008
Cited alongside, same era.
Discrete Choice Methods with Simulation
Train, Keneth · 2009
Cited alongside, same era.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yue, Yisong and Joachims, T · 2009
Cited alongside, same era.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Agarwal, Alekh, Dekel, Ofer, and Xiao, Lin · 2010
Cited alongside, same era.
From bandits to experts: On the value of side-observations
Mannor, Shie and Shamir, Ohad · 2011
Cited alongside, same era.
Active learning using smooth relative regret approximations with applications
Ailon, Nir, Begleiter, Ron, and Ezra, Esther · 2012
Later among the works it cites.
Toward a classification of finite partial-monitoring games
Antos, András, Bartók, Gábor, Pál, Dávid, and Szepesvári, Csaba · 2012
Later among the works it cites.
Decoupling exploration and exploitation in multi-armed bandits
Avner, Orly, Mannor, Shie, and Shamir, Ohad · 2012
Later among the works it cites.
An adaptive algorithm for finite stochastic partial monitoring
Bartók, Gábor, Zolghadr, Navid, and Szepesvári, Csaba · 2012
Later among the works it cites.
Large-scale validation and analysis of interleaved search evaluation
Chapelle, O., Joachims, T., Radlinski, F., and Yue, Yisong · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beat the mean bandit
Yue, Yisong and Joachims, Thorsten · 2011
Cited alongside, same era.
Yue, Yisong, Broder, Josef, Kleinberg, Robert, and Joachims, Thorsten · 2012
Later among the works it cites.