Fetching the paper…
Reading the bibliography…
Query performance prediction (QPP) aims to estimate the retrieval quality of a search system for a query without human relevance judgments.
Language Models are Few-Shot Learners. In NeurIPS . 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Hierarchical Dependence-aware Evaluation Measures for Conversational Search. In SIGIR . 1935–1939
Guglielmo Faggioli, Marco Ferrante, Nicola Ferro, Raffaele Perego, and Nicola Tonellotto. 2021a · 1939
Earlier work this paper cites.
Large Language Models Can Accurately Predict Searcher Preferences. In SIGIR . 1930–1940
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 1940
Earlier work this paper cites.
Are Large Language Models Good at Utility Judgments?. In SIGIR . 1941–1951
Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024 · 1951
Earlier work this paper cites.
Cast-19: A Dataset for Conversational Information Seeking. In SIGIR . 1985–1988
Jeffrey Dalton, Chenyan Xiong, Vaibhav Kumar, and Jamie Callan. 2020b · 1988
Earlier work this paper cites.
Okapi at TREC-3
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al · 1995
Earlier work this paper cites.
Document Language Models, Query Models, and Risk Minimization for Information Retrieval. In SIGIR . 111–119
John Lafferty and Chengxiang Zhai. 2001 · 2001
Earlier work this paper cites.
Relevance-Based Language Models. In SIGIR . 120–127
Victor Lavrenko and W. Bruce Croft. 2001 · 2001
Earlier work this paper cites.
Ranking Retrieval Systems without Relevance Judgments. In SIGIR . 66–73
Ian Soboroff, Charles Nicholas, and Patrick Cahan. 2001 · 2001
Earlier work this paper cites.
Predicting Query Performance. In SIGIR . 299–306
Steve Cronen-Townsend, Yun Zhou, and W. Bruce Croft. 2002 · 2002
Earlier work this paper cites.
Cumulated Gain-Based Evaluation of IR Techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
Automatic Ranking of Retrieval Systems in Imperfect Environments. In SIGIR . 379–380
Rabia Nuray and Fazli Can. 2003 · 2003
Earlier work this paper cites.
Query Difficulty, Robustness, and Selective Application of Query Expansion. In ECIR . 127–137
Giambattista Amati, Claudio Carpineto, and Giovanni Romano. 2004 · 2004
Earlier work this paper cites.
Learning to Rank System Configurations. In CIKM . 2001–2004
Romain Deveaud, Josiane Mothe, and Jian-Yun Nie. 2016 · 2004
Earlier work this paper cites.
Automatic Ranking of Information Retrieval Systems using Data Fusion
Rabia Nuray and Fazli Can. 2006 · 2006
Earlier work this paper cites.
Ranking Robustness: A Novel Framework to Predict Query Performance. In CIKM . 567–574
Yun Zhou and W. Bruce Croft. 2006 · 2006
Earlier work this paper cites.
Query Hardness Estimation Using Jensen-Shannon Divergence Among Multiple Scoring Functions. In ECIR . 198–209
Javed A Aslam and Virgil Pavlu. 2007 · 2007
Earlier work this paper cites.
Performance Prediction Using Spatial Autocorrelation. In SIGIR . 583–590
Fernando Diaz. 2007 · 2007
Earlier work this paper cites.
Overview of the TREC 2007 Legal Track.. In TREC
Stephen Tomlinson, Douglas W Oard, Jason R Baron, and Paul Thompson. 2007 · 2007
Earlier work this paper cites.
Query Performance Prediction in Web Search Environments. In SIGIR . 543–550
Yun Zhou and W. Bruce Croft. 2007 · 2007
Earlier work this paper cites.
A Survey of Pre-retrieval Query Performance Predictors. In CIKM . 1419–1420
Claudia Hauff, Djoerd Hiemstra, and Franciska de Jong. 2008 · 2008
Earlier work this paper cites.
Estimating the Query Difficulty for Information Retrieval
David Carmel and Elad Yom-Tov. 2010 · 2010
Earlier work this paper cites.
Standard Deviation as a Query Hardness Estimator. In SPIRE . 207–212
Joaquín Pérez-Iglesias and Lourdes Araujo. 2010 · 2010
Earlier work this paper cites.
Using Statistical Decision Theory and Relevance Models for Query-performance Prediction. In SIGIR . 259–266
Anna Shtok, Oren Kurland, and David Carmel. 2010 · 2010
Earlier work this paper cites.
Improved Query Performance Prediction Using Standard Deviation. In SIGIR . 1089–1090
Ronan Cummins, Joemon Jose, and Colm O’Riordan. 2011 · 2011
Earlier work this paper cites.
On the Usefulness of Query Features for Learning to Rank. In CIKM . 2559–2562
Craig Macdonald, Rodrygo LT Santos, and Iadh Ounis. 2012 · 2012
Earlier work this paper cites.
Predicting Query Performance by Query-Drift Estimation
Anna Shtok, Oren Kurland, David Carmel, Fiana Raiber, and Gad Markovits. 2012 · 2012
Earlier work this paper cites.
Efficient and Effective Retrieval using Selective Pruning. In WSDM . 63–72
Nicola Tonellotto, Craig Macdonald, and Iadh Ounis. 2013 · 2013
Earlier work this paper cites.
Query Performance Prediction by Considering Score Magnitude and Variance Together. In CIKM . 1891–1894
Yongquan Tao and Shengli Wu. 2014 · 2014
Earlier work this paper cites.
Features of Disagreement Between Retrieval Effectiveness Measures. In SIGIR . 847–850
Timothy Jones, Paul Thomas, Falk Scholer, and Mark Sanderson. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In ICLR
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Ranking Retrieval Systems using Pseudo Relevance Judgments
Sri Devi Ravana, Prabha Rajagopal, and Vimala Balakrishnan. 2015 · 2015
Earlier work this paper cites.
The Effect of Pooling and Evaluation Depth on IR Metrics
Xiaolu Lu, Alistair Moffat, and J. Shane Culpepper. 2016 · 2016
Earlier work this paper cites.
Towards Automatic Generation of Relevance Judgments for a Test Collection. In ICDIM . IEEE, 121–126
Mireille Makary, Michael Oakes, and Fadi Yamout. 2016 · 2016
Earlier work this paper cites.
Using Supervised Machine Learning to Automatically Build Relevance Judgments for a Test Collection. In 2017 28th International Workshop on Database and Expert Systems Applications (DEXA) . IEEE, 108–112
Mireille Makary, Michael Oakes, Ruslan Mitkov, and Fadi Yammout. 2017 · 2017
Earlier work this paper cites.
Computing Maximized Effectiveness Distance for Recall-based Metrics
Alistair Moffat. 2017 · 2017
Earlier work this paper cites.
An Enhanced Approach to Query Performance Prediction Using Reference Lists. In SIGIR . 869–872
Haggai Roitman. 2017 · 2017
Earlier work this paper cites.
Tasks, Queries, and Rankers in Pre-Retrieval Performance Prediction. In Proceedings of the 22nd Australasian Document Computing Symposium . 1–4
Paul Thomas, Falk Scholer, Peter Bailey, and Alistair Moffat. 2017 · 2017
Earlier work this paper cites.
Query Performance Prediction and Effectiveness Evaluation Without Relevance Judgments: Two Sides of the Same Coin. In SIGIR . 1233–1236
Stefano Mizzaro, Josiane Mothe, Kevin Roitero, and Md Zia Ullah. 2018 · 2018
Earlier work this paper cites.
Query Variation Performance Prediction for Systematic Reviews. In SIGIR . 1089–1092
Harrisen Scells, Leif Azzopardi, Guido Zuccon, and Bevan Koopman. 2018 · 2018
Earlier work this paper cites.
Neural Query Performance Prediction Using Weak Supervision from Multiple Signals. In SIGIR . 105–114
Hamed Zamani, W. Bruce Croft, and J. Shane Culpepper. 2018 · 2018
Earlier work this paper cites.
Overview of the TREC 2019 Deep Learning Track. In TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Correlation, Prediction and Ranking of Evaluation Metrics in Information Retrieval. In ECIR . 636–651
Soumyajit Gupta, Mucahid Kutlu, Vivek Khetan, and Matthew Lease. 2019 · 2019
Earlier work this paper cites.
Performance Prediction for Non-factoid Question Answering. In ICTIR . 55–58
Helia Hashemi, Hamed Zamani, and W. Bruce Croft. 2019 · 2019
Cited alongside, same era.
Overview of the TREC 2020 Deep Learning Track. In TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2020 · 2020
Cited alongside, same era.
CAsT 2020: The Conversational Assistance Track Overview. In Text Retrieval Conference
Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020a · 2020
Cited alongside, same era.
Overview of the TREC 2021 Deep Learning Track. In TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Fernando Campos, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
A Study of a Gain Based Approach for Query Aspects in Recall Oriented Tasks
Giorgio Maria Di Nunzio and Guglielmo Faggioli. 2021 · 2021
Cited alongside, same era.
Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling. In SIGIR . 113–122
Zero-Shot Listwise Document Reranking with a Large Language Model
Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023b · 2023
Later among the works it cites.
One-Shot Labeling for Automatic Relevance Estimation. In SIGIR . 2230–2235
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Later among the works it cites.
Performance Prediction for Conversational Search Using Perplexities of Query Rewrites. In QPP++2023 . 25–28
Chuan Meng, Mohammad Aliannejadi, and Maarten de Rijke. 2023a · 2023
Later among the works it cites.
IQPP: A Benchmark for Image Query Performance Prediction. In SIGIR . 2953–2963
Eduard Poesina, Radu Tudor Ionescu, and Josiane Mothe. 2023 · 2023
Later among the works it cites.
RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023a · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021 · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models. In ICLR
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Cited alongside, same era.
Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations. In SIGIR . 2356–2362
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021 · 2021
Cited alongside, same era.
LeCaRD: A Legal Case Retrieval Dataset for Chinese Law System. In SIGIR . 2342–2348
Yixiao Ma, Yunqiu Shao, Yueyue Wu, Yiqun Liu, Ruizhe Zhang, Min Zhang, and Shaoping Ma. 2021 · 2021
Cited alongside, same era.
The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
Conversations Powered by Cross-Lingual Knowledge. In SIGIR . 1442–1451
Weiwei Sun, Chuan Meng, Qi Meng, Zhaochun Ren, Pengjie Ren, Zhumin Chen, and Maarten de Rijke. 2021 · 2021
Cited alongside, same era.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In ICLR
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023b · 2023
Later among the works it cites.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, et al · 2023
Later among the works it cites.
Performance Prediction for Multi-hop Questions
Mohammadreza Samadi and Davood Rafiei. 2023 · 2023
Later among the works it cites.
Camoscio: An Italian Instruction-tuned Llama
Andrea Santilli and Emanuele Rodolà. 2023 · 2023
Later among the works it cites.
Unsupervised Query Performance Prediction for Neural Models utilising Pairwise Rank Preferences. In SIGIR . 2486–2490
Ashutosh Singh, Debasis Ganguly, Suchana Datta, and Craig McDonald. 2023 · 2023
Later among the works it cites.
Evaluating the Zero-shot Robustness of Instruction-tuned Language Models
Jiuding Sun, Chantal Shaib, and Byron C. Wallace. 2023a · 2023
Later among the works it cites.
Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models
Raphael Tang, Xinyu Zhang, Xueguang Ma, Jimmy Lin, and Ferhan Ture. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
On Coherence-based Predictors for Dense Query Performance Prediction
Maria Vlachou and Craig Macdonald. 2023 · 2023
Later among the works it cites.
Mingxue Xu, Yao Lei Xu, and Danilo P Mandic. 2023 · 2023
Later among the works it cites.
Entropy-Based Query Performance Prediction for Neural Information Retrieval Systems. In QPP++2023 . 37–44
Oleg Zendel, Binsheng Liu, J. Shane Culpepper, and Falk Scholer. 2023 · 2023
Later among the works it cites.
Instruction Tuning for Large Language Models: A Survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al · 2023
Later among the works it cites.
Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models
Xinyu Zhang, Sebastian Hofstätter, Patrick Lewis, Raphael Tang, and Jimmy Lin. 2023c · 2023
Later among the works it cites.
Yue Zhang, Leyang Cui, Deng Cai, Xinting Huang, Tao Fang, and Wei Bi. 2023a · 2023
Later among the works it cites.
Large Language Models for Information Retrieval: A Survey
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen. 2023 · 2023
Later among the works it cites.
Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels
Honglei Zhuang, Zhen Qin, Kai Hui, Junru Wu, Le Yan, Xuanhui Wang, and Michael Berdersky. 2023b · 2023
Later among the works it cites.
Open-source Large Language Models are Strong Zero-shot Query Likelihood Models for Document Ranking
Shengyao Zhuang, Bing Liu, Bevan Koopman, and Guido Zuccon. 2023a · 2023
Later among the works it cites.
A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models
Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. 2023c · 2023
Later among the works it cites.
Can We Use Large Language Models to Fill Relevance Judgment Holes?
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi. 2024 · 2024
Closest in time.
Generative Retrieval with Few-shot Indexing
Arian Askari, Chuan Meng, Mohammad Aliannejadi, Zhaochun Ren, Evangelos Kanoulas, and Suzan Verberne. 2024 · 2024
Closest in time.
Nuo Chen, Jiqun Liu, Xiaoyu Dong, Qijiong Liu, Tetsuya Sakai, and Xiao-Ming Wu. 2024 · 2024
Closest in time.
Leveraging LLMs for Unsupervised Dense Retriever Ranking
Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, and Guido Zuccon. 2024 · 2024
Closest in time.
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
Shengjie Ma, Chong Chen, Qi Chu, and Jiaxin Mao. 2024 · 2024
Closest in time.
Query Performance Prediction for Conversational Search and Beyond. In SIGIR
Chuan Meng. 2024 · 2024
Closest in time.
Ranked List Truncation for Large Language Model-based Re-Ranking. In SIGIR . 141–151
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024 · 2024
Closest in time.
A Survey of Conversational Search
Fengran Mo, Kelong Mao, Ziliang Zhao, Hongjin Qian, Haonan Chen, Yiruo Cheng, Xiaoxi Li, Yutao Zhu, Zhicheng Dou, and Jian-Yun Nie. 2024b · 2024
Closest in time.
Evaluating Retrieval Quality in Retrieval-Augmented Generation
Alireza Salemi and Hamed Zamani. 2024 · 2024
Closest in time.
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
Rikiya Takehi, Ellen M Voorhees, and Tetsuya Sakai. 2024 · 2024
Closest in time.
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
Shivani Upadhyay, Ehsan Kamalloo, and Jimmy Lin. 2024a · 2024
Closest in time.
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, and Jimmy Lin. 2024b · 2024
Closest in time.
Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing
Le Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman, Xuanhui Wang, Michael Bendersky, and Harrie Oosterhuis. 2024 · 2024
Closest in time.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Improving the Reusability of Conversational Search Test Collections. In ECIR . 196–213
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi. 2025 · 2025
Closest in time.
Query Performance Prediction: Theory, Techniques and Applications. In WSDM . 991–994
Negar Arabzadeh, Chuan Meng, Mohammad Aliannejadi, and Ebrahim Bagheri. 2025 · 2025
Closest in time.
Self-seeding and Multi-intent Self-instructing LLMs for Generating Intent-aware Information-Seeking Dialogs. In NAACL
Arian Askari, Roxana Petcu, Chuan Meng, Mohammad Aliannejadi, Amin Abolghasemi, Evangelos Kanoulas, and Suzan Verberne. 2025 · 2025
Closest in time.
Zero-Shot and Efficient Clarification Need Prediction in Conversational Search. In ECIR . 389–404
Lili Lu, Chuan Meng, Federico Ravenda, Mohammad Aliannejadi, and Fabio Crestani. 2025 · 2025
Closest in time.
QPP++ 2025: Query Performance Prediction and its Applications in the Era of Large Language Models. In ECIR . 319–325
Chuan Meng, Guglielmo Faggioli, Mohammad Aliannejadi, Nicola Ferro, and Josiane Mothe. 2025a · 2025
Closest in time.
Conversational Search: From Fundamentals to Frontiers in the LLM Era. In SIGIR
Fengran Mo, Chuan Meng, Mohammad Aliannejadi, and Jian-Yun Nie. 2025 · 2025
Closest in time.