Fetching the paper…
Reading the bibliography…
Assessors make preference judgments faster and more consistently than graded judgments.
The simple scalability of documents
Mark E. Rorvig. 1990 · 1990
Earlier work this paper cites.
Determining the effectiveness of retrieval algorithms
H. P. Frei and P. Schäuble. 1991 · 1991
Earlier work this paper cites.
Measuring retrieval effectiveness based on user preference of documents
Y. Y. Yao. 1995 · 1995
Earlier work this paper cites.
Variations in relevance judgments and the measurement of retrieval effectiveness. In 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . Melbourne, Australia, 315–323
Ellen M. Voorhees. 1998 · 1998
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
Eye-Tracking Analysis of User Behavior in WWW Search. In 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . 478–479
Laura A. Granka, Thorsten Joachims, and Geri Gay. 2004 · 2004
Earlier work this paper cites.
Evaluating evaluation metrics based on the bootstrap. In 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . Seattle, Washington, 525–532
Tetsuya Sakai. 2006 · 2006
Earlier work this paper cites.
Crowdsourcing for relevance evaluation
Omar Alonso, Daniel E. Rose, and Benjamin Stewart. 2008 · 2008
Earlier work this paper cites.
Relevance assessment: Are judges exchangeable and does it matter. In 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . Singapore, 667–674
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P. de Vries, and Emine Yilmaz. 2008 · 2008
Earlier work this paper cites.
A test collection of preference judgments. In SIGIR 2008 Workshop on Beyond Binary Relevance: Preferences, Diversity, and Set-Level Judgments . Singapore
Ben Carterette, Paul Bennett, and Olivier Chapelle. 2008a · 2008
Earlier work this paper cites.
Evaluation measures for preference judgments. In 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . Singapore, 685–686
Ben Carterette and Paul N. Bennett. 2008 · 2008
Earlier work this paper cites.
Expected reciprocal rank for graded relevance. In 18th ACM Conference on Information and Knowledge Management . Hong Kong, China, 621–630
Olivier Chapelle, Donald Metlzer, Ya Zhang, and Pierre Grinspan. 2009 · 2009
Earlier work this paper cites.
From RankNet to LambdaRank to LambdaMART: An overview
Christopher J. C. Burges. 2010 · 2010
Earlier work this paper cites.
A similarity measure for indefinite rankings
William Webber, Alistair Moffat, and Justin Zobel. 2010 · 2010
Cited alongside, same era.
An analysis of assessor behavior in crowdsourced preference judgments. In SIGIR 2010 Workshop on Crowdsourcing for Search Evaluation
Dongqing Zhu and Ben Carterette. 2010 · 2010
Cited alongside, same era.
Ranking from pairs and triplets: Information quality, evaluation methods and query complexity. In 4th ACM International Conference on Web Search and Data Mining . Hong Kong, China, 105–114
Kira Radinsky and Nir Ailon. 2011 · 2011
Cited alongside, same era.
Using preference judgments for novel document retrieval. In 35th International ACM SIGIR Conference on Research and Development in Information Retrieval . Portland, Oregon, 861–870
Praveen Chandar and Ben Carterette. 2012 · 2012
Cited alongside, same era.
Crowdsourcing for information retrieval
Matthew Lease and Emine Yilmaz. 2012 · 2012
Cited alongside, same era.
On crowdsourcing relevance magnitudes for information retrieval evaluation
Eddy Maddalena, Stefano Mizzaro, Falk Scholer, and Andrew Turpin. 2017 · 2017
Later among the works it cites.
The Notion of Relevance in Information Science: Everybody knows what relevance is. But, what is it really?
Tefko Saracevic. 2017 · 2017
Later among the works it cites.
Eliciting pairwise preferences in recommender systems. In 12th ACM Conference on Recommender Systems . Vancouver, British Columbia, 329–337
Saikishore Kalloori, Francesco Ricci, and Rosella Gennari. 2018 · 2018
Later among the works it cites.
Pairwise crowd judgments: Preference, absolute, and ratio. In 23rd Australasian Document Computing Symposium . Dunedin, New Zealand
Ziying Yang, Alistair Moffat, and Andrew Turpin. 2018 · 2018
Later among the works it cites.
Patterns of Search Result Examination: Query to First Action. In 28th ACM International Conference on Information and Knowledge Management . 1833–1842
Mustafa Abualsaud and Mark D. Smucker. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Top-k learning to rank: Labeling, ranking and evaluation. In 35th International ACM SIGIR Conference on Research and Development in Information Retrieval . Portland, Oregon, 751–760
Shuzi Niu, Jiafeng Guo, Yanyan Lan, and Xueqi Cheng. 2012 · 2012
Cited alongside, same era.
A document rating system for preference judgements. In 36th International ACM SIGIR Conference on Research and Development in Information Retrieval . Dublin, Ireland, 909–912
Maryam Bashir, Jesse Anderton, Jie Wu, Peter B Golbus, Virgil Pavlu, and Javed A. Aslam. 2013 · 2013
Cited alongside, same era.
Preference based evaluation measures for novelty and diversity. In 36th International ACM SIGIR Conference on Research and Development in Information Retrieval . Dublin, Ireland, 413–422
Praveen Chandar and Ben Carterette. 2013 · 2013
Cited alongside, same era.
Pairwise ranking aggregation in a crowdsourced setting. In 6th ACM International Conference on Web Search and Data Mining . Rome, Italy, 193–202
Xi Chen, Paul N. Bennett, Kevyn Collins-Thompson, and Eric Horvitz. 2013 · 2013
Cited alongside, same era.
User Intent and Assessor Disagreement in Web Search Evaluation. In 22nd ACM International Conference on Information and Knowledge Management . San Francisco, California, 699–708
Gabriella Kazai, Emine Yilmaz, Nick Craswell, and S.M.M. Tahaghoghi. 2013 · 2013
Cited alongside, same era.
Relevance dimensions in preference-based IR evaluation. In 36th International ACM SIGIR Conference on Research and Development in Information Retrieval . Dublin, Ireland, 913–916
Jinyoung Kim, Gabriella Kazai, and Imed Zitouni. 2013 · 2013
Cited alongside, same era.
Is top- k k sufficient for ranking?. In 22nd ACM International Conference on Information and Knowledge Management . San Francisco, California, 1261–1270
Yanyan Lan, Shuzi Niu, Jiafeng Guo, and Xueqi Cheng. 2013 · 2013
Cited alongside, same era.
Later among the works it cites.
WaterlooClarke at the TREC 2019 Conversational Assistant Track. In 28th Text REtrieval Conference . Gaithersburg, Maryland
Charles L. A. Clarke. 2019 · 2019
Later among the works it cites.
CAsT 2019: The Conversational Assistance Track overview. In 28th Text REtrieval Conference . Gaithersburg, Maryland
Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding. In Annual Conference of the North American Chapter of the Association for Computational Linguistics . Minneapolis, Minnesota
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Evaluating preference collection methods for interactive ranking analytics. In 2019 CHI Conference on Human Factors in Computing Systems . Glasgow, Scotland Uk, Article Paper 512, 11 pages
Caitlin Kuhlman, Diana Doherty, Malika Nurbekova, Goutham Deva, Zarni Phyo, Paul-Henry Schoenhagen, MaryAnn VanValkenburg, Elke Rundensteiner, and Lane Harrison. 2019 · 2019
Later among the works it cites.
Good evaluation measures based on document preferences. In 43st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . Xi’an, China
Tetsuya Sakai and Zhaohao Zeng. 2020 · 2020
Closest in time.
Preference-based Evaluation Metrics for Web Image Search. In 43st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval . Xi’an, China
Xiaohui Xie, Jiaxin Mao, Yiqun Liu, Maarten de Rijke, Haitian Chen, Min Zhang, and Shaoping Ma. 2020 · 2020
Closest in time.
Statistical consistency of top- k k ranking. In 22nd International Conference on Neural Information Processing Systems . Vancouver, British Columbia, 2098–2106
Fen Xia, Tie-Yan Liu, and Hang Li. 2009 · 2098
Closest in time.