Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of safe online learning to re-rank, where user feedback is used to improve the quality of displayed lists.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
K. Järvelin and J. Kekäläinen · 2002
Earlier work this paper cites.
Optimizing search engines using clickthrough data
T. Joachims · 2002
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Minimally invasive randomization for collecting unbiased preferences from clickthrough logs
F. Radlinski and T. Joachims · 2006
Earlier work this paper cites.
Predicting clicks: Estimating the click-through rate for new ads
M. Richardson, E. Dominowska, and R. Ragno · 2007
Earlier work this paper cites.
An experimental comparison of click position-bias models
N. Craswell, O. Zoeter, M. Taylor, and B. Ramsey · 2008
Earlier work this paper cites.
Learning diverse rankings with multi-armed bandits
F. Radlinski, R. Kleinberg, and T. Joachims · 2008
Earlier work this paper cites.
Introduction to Algorithms, 3rd Edition
T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein · 2009
Earlier work this paper cites.
Efficient multiple-click models in web search
F. Guo, C. Liu, and Y. M. Wang · 2009
Earlier work this paper cites.
Learning to rank for information retrieval
T.-Y. Liu · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Online learning for recency search ranking using real-time user feedback
T. Moon, L. Li, W. Chu, C. Liao, Z. Zheng, and Y. Chang · 2010
Cited alongside, same era.
Letor: A benchmark collection for research on learning to rank for information retrieval
T. Qin, T.-Y. Liu, J. Xu, and H. Li · 2010
Cited alongside, same era.
Learning from logged implicit exploration data
A. Strehl, J. Langford, L. Li, and S. M. Kakade · 2010
Cited alongside, same era.
Linear submodular bandits and their application to diversified retrieval
Y. Yue and C. Guestrin · 2011
Cited alongside, same era.
Reusing historical interaction data for faster online learning to rank for ir
K. Hofmann, A. Schuth, S. Whiteson, and M. de Rijke · 2013
Cited alongside, same era.
Ad click prediction: a view from the trenches
H. B. McMahan, G. Holt, D. Sculley, M. Young, D. Ebner, J. Grady, L. Nie, T. Phillips, E. Davydov, D. Golovin, et al · 2013
Conservative bandits
Y. Wu, R. Shariff, T. Lattimore, and C. Szepesvári · 2016
Later among the works it cites.
Click-based hot fixes for underperforming torso queries
M. Zoghi, T. Tunys, L. Li, D. Jose, J. Chen, C. M. Chin, and M. de Rijke · 2016
Later among the works it cites.
Cascading bandits for large-scale recommendation problems
S. Zong, H. Ni, K. Sung, N. R. Ke, Z. Wen, and B. Kveton · 2016
Later among the works it cites.
TREC 2017 common core track overview
J. Allan, D. Harman, E. Kanoulas, D. Li, C. Van Gysel, and E. Voorhees · 2017
Later among the works it cites.
Efficient cost-aware cascade ranking in multi-stage retrieval
R.-C. Chen, L. Gallagher, R. Blanco, and J. S. Culpepper · 2017
Later among the works it cites.
On application of learning to rank for e-commerce search
S. K. Karmaker Santu, P. Sondhi, and C. Zhai · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ranked bandits in metric spaces: Learning diverse rankings over large document collections
A. Slivkins, F. Radlinski, and S. Gollapudi · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Cited alongside, same era.
Click Models for Web Search
A. Chuklin, I. Markov, and M. de Rijke · 2015
Cited alongside, same era.
Learning to rank: Regret lower bounds and efficient algorithms
R. Combes, S. Magureanu, A. Proutiere, and C. Laroche · 2015
Cited alongside, same era.
Gathering additional feedback on search results by multi-armed bandits with respect to production ranking
A. Vorobev, D. Lefortier, G. Gusev, and P. Serdyukov · 2015
Cited alongside, same era.
DCM bandits: Learning to rank with multiple clicks
S. Katariya, B. Kveton, C. Szepesvari, and Z. Wen · 2016
Cited alongside, same era.
Later among the works it cites.
Conservative contextual linear bandits
A. Kazerouni, M. Ghavamzadeh, Y. Abbasi, and B. Van Roy · 2017
Later among the works it cites.
Cascade ranking for operational e-commerce search
S. Liu, F. Xiao, W. Ou, and L. Si · 2017
Later among the works it cites.
Online learning to rank in stochastic click models
M. Zoghi, T. Tunys, M. Ghavamzadeh, B. Kveton, C. Szepesvari, and Z. Wen · 2017
Later among the works it cites.
Toprank: A practical algorithm for online stochastic ranking
T. Lattimore, B. Kveton, S. Li, and C. Szepesvari · 2018
Closest in time.
Position bias estimation for unbiased learning to rank in personal search
X. Wang, N. Golbandi, M. Bendersky, D. Metzler, and M. Najork · 2018
Closest in time.
A practical deep online ranking system in e-commerce recommendation
Y. Yan, Z. Liu, M. Zhao, W. Guo, W. P. Yan, and Y. Bao · 2018
Closest in time.