Fetching the paper…
Reading the bibliography…
In search settings, calibrating the scores during the ranking process to quantities such as click-through rates or relevance levels enhances a system's usefulness and trustworthiness for downstream users.
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Document Expansion by Query Prediction
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019b · 1904
Earlier work this paper cites.
Multi-Stage Document Ranking with BERT
Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019a · 1910
Earlier work this paper cites.
Yahoo! Learning to Rank Challenge Overview
Olivier Chapelle and Yi Chang. 2011 · 1938
Earlier work this paper cites.
Reliability of Subjective Probability Forecasts of Precipitation and Temperature
Allan H. Murphy and Robert L. Winkler. 1977 · 1977
Earlier work this paper cites.
Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods
John Platt. 2000 · 2000
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Bianca Zadrozny and Charles Elkan. 2002 · 2002
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020 · 2003
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Query performance prediction in web search environments
Yun Zhou and W. Bruce Croft. 2007 · 2007
Earlier work this paper cites.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y.-Lan Boureau. 2022 · 2012
Earlier work this paper cites.
Predicting Query Performance by Query-Drift Estimation
Anna Shtok, Oren Kurland, David Carmel, Fiana Raiber, and Gad Markovits. 2012 · 2012
Earlier work this paper cites.
Time-based calibration of effectiveness measures
Mark D. Smucker and Charles L.A. Clarke. 2012 · 2012
Earlier work this paper cites.
Introducing LETOR 4.0 Datasets
Tao Qin and Tie-Yan Liu. 2013 · 2013
Earlier work this paper cites.
CTR prediction for contextual advertising: learning-to-rank approach
Yukihiro Tagami, Shingo Ono, Koji Yamamoto, Koji Tsukamoto, and Akira Tajima. 2013 · 2013
Earlier work this paper cites.
Ranking and Calibrating Click-Attributed Purchases in Performance Display Advertising
Sougata Chaudhuri, Abraham Bagherjeiran, and James Liu. 2017 · 2017
Earlier work this paper cites.
Fast Ranking with Additive Ensembles of Oblivious and Non-Oblivious Regression Trees
Domenico Dato, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando, Raffaele Perego, Nicola Tonellotto, and Rossano Venturini. 2017 · 2017
Earlier work this paper cites.
On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018 · 2018
Earlier work this paper cites.
e-SNLI: Natural Language Inference with Natural Language Explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Overview of the NTCIR-14 We Want Web Task
Jiaxin Mao, Tetsuya Sakai, Cheng Luo, Peng Xiao, Yiqun Liu, and Zhicheng Dou. 2019 · 2019
Cited alongside, same era.
Choppy: Cut Transformer for Ranked List Truncation
Dara Bahri, Yi Tay, Che Zheng, Donald Metzler, and Andrew Tomkins. 2020 · 2020
Cited alongside, same era.
Regression Compatible Listwise Objectives for Calibrated Ranking with Binary Relevance
Aijun Bai, Rolf Jagerman, Zhen Qin, Le Yan, Pratyush Kar, Bing-Rong Lin, Xuanhui Wang, Michael Bendersky, and Marc Najork. 2023 · 2023
Later among the works it cites.
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, and Kathleen McKeown. 2023 · 2023
Later among the works it cites.
Perspectives on Large Language Models for Relevance Judgment
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023a · 2023
Later among the works it cites.
ExaRanker: Explanation-Augmented Neural Ranker
Fernando Ferraretto, Thiago Laitz, Roberto Lotufo, and Rodrigo Nogueira. 2023 · 2023
Later among the works it cites.
Predictive Uncertainty-based Bias Mitigation in Ranking
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Omar Khattab and Matei Zaharia. 2020 · 2020
Cited alongside, same era.
Generate Neural Template Explanations for Recommendation
Lei Li, Yongfeng Zhang, and Li Chen. 2020 · 2020
Cited alongside, same era.
Document Ranking with a Pretrained Sequence-to-Sequence Model
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020 · 2020
Cited alongside, same era.
Daniel Cohen, Bhaskar Mitra, Oleg Lesota, Navid Rekabsaz, and Carsten Eickhoff. 2021 · 2021
Cited alongside, same era.
TREC Deep Learning Track: Reusable Test Collections in the Large Data Regime
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Ellen M. Voorhees, and Ian Soboroff. 2021 · 2021
Cited alongside, same era.
On the Calibration and Uncertainty of Neural Learning to Rank Models for Conversational Search
Gustavo Penha and Claudia Hauff. 2021 · 2021
Cited alongside, same era.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021 · 2021
Cited alongside, same era.
Maria Heuss, Daniel Cohen, Masoud Mansoury, Maarten de Rijke, and Carsten Eickhoff. 2023 · 2023
Later among the works it cites.
Query Performance Prediction: From Ad-hoc to Conversational Search
Chuan Meng, Negar Arabzadeh, Mohammad Aliannejadi, and Maarten De Rijke. 2023 · 2023
Later among the works it cites.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2023 · 2023
Later among the works it cites.
Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S. Jaakkola, and Regina Barzilay. 2023 · 2023
Later among the works it cites.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2023 · 2023
Later among the works it cites.
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. 2023 · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023 · 2023
Later among the works it cites.
Using Natural Language Explanations to Rescale Human Judgments
Manya Wadhwa, Jifan Chen, Junyi Jessy Li, and Greg Durrett. 2023 · 2023
Later among the works it cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Later among the works it cites.
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Later among the works it cites.
Query Performance Prediction: From Fundamentals to Advanced Techniques
Negar Arabzadeh, Chuan Meng, Mohammad Aliannejadi, and Ebrahim Bagheri. 2024 · 2024
Closest in time.
InRanker: Distilled Rankers for Zero-shot Information Retrieval
Thiago Laitz, Konstantinos Papakostas, Roberto Lotufo, and Rodrigo Nogueira. 2024 · 2024
Closest in time.
A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models
Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. 2024 · 2024
Closest in time.