Fetching the paper…
Reading the bibliography…
We investigate the usefulness of generative Large Language Models (LLMs) in generating training data for cross-encoder re-rankers in a novel direction: generating synthetic documents instead of synthetic queries.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Www’18 open challenge: financial opinion mining and question answering. In Companion proceedings of the the web conference 2018 . 1941–1942
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018 · 1942
Earlier work this paper cites.
Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval. In SIGIR’94 . Springer, 232–241
Stephen E Robertson and Steve Walker. 1994 · 1994
Earlier work this paper cites.
The history of information retrieval research
Mark Sanderson and W Bruce Croft. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset. In CoCo@ NIPs
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering. In Proceedings of the 2015 conference on empirical methods in natural language processing . 2013–2018
Yi Yang, Wen-tau Yih, and Christopher Meek. 2015 · 2018
Earlier work this paper cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhilu Zhang and Mert Sabuncu. 2018 · 2018
Earlier work this paper cites.
ELI5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Meddialog: a large-scale medical dialogue dataset
Shu Chen, Zeqian Ju, Xiangyu Dong, Hongchao Fang, Sicheng Wang, Yue Yang, Jiaqi Zeng, Ruisi Zhang, Ruoyu Zhang, Meng Zhou, et al · 2020
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees. 2020 · 2020
Earlier work this paper cites.
Improving efficient neural ranking models with cross-architecture knowledge distillation
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020 · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020 . Association for Computational Linguistics, Online, 4163–4174
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Cited alongside, same era.
Expansion via prediction of importance with contextualization. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 1573–1576
Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto, Nazli Goharian, and Ophir Frieder. 2020 · 2020
Cited alongside, same era.
Document ranking with a pretrained sequence-to-sequence model
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. 2020 · 2020
Evaluating the feasibility of ChatGPT in healthcare: an analysis of multiple clinical and research scenarios
Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami. 2023 · 2023
Closest in time.
Elasticsearch
BV Elasticsearch. 2023 · 2023
Closest in time.
Perspectives on Large Language Models for Relevance Judgment
Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, et al · 2023
Closest in time.
Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova. 2023 · 2023
Closest in time.
How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Overview of the TREC 2020 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021 · 2021
Cited alongside, same era.
Pretrained transformers for text ranking: Bert and beyond
Jimmy Lin, Rodrigo Nogueira, and Andrew Yates. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Deep query likelihood model for information retrieval. In European Conference on Information Retrieval . Springer, 463–470
Shengyao Zhuang, Hang Li, and Guido Zuccon. 2021 · 2021
Cited alongside, same era.
Shengyao Zhuang and Guido Zuccon. 2021 · 2021
Cited alongside, same era.
Inpars: Data augmentation for information retrieval using large language models
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Cited alongside, same era.
Promptagator: Few-shot dense retrieval from 8 examples
Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B Hall, and Ming-Wei Chang. 2022 · 2022
Cited alongside, same era.
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023 · 2023
Closest in time.
InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval
Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023 · 2023
Closest in time.
Explain like I am BM25: Interpreting a Dense Model’s Ranked-List with a Sparse Approximation
Michael Llordes, Debasis Ganguly, Sumit Bhatia, and Chirag Agarwal. 2023 · 2023
Closest in time.
Towards Making the Most of ChatGPT for Machine Translation
Keqin Peng, Liang Ding, Qihuang Zhong, Li Shen, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2023 · 2023
Closest in time.
ChatGPT applications in medical, dental, pharmacy, and public health education: A descriptive study highlighting the advantages and limitations
Malik Sallam, Nesreen Salim, Muna Barakat, and Alaa Al-Tammemi. 2023 · 2023
Closest in time.
Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agent
Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023 · 2023
Closest in time.
A short survey of viewing large language models in legal aspect
Zhongxiang Sun. 2023 · 2023
Closest in time.
Applying BERT and ChatGPT for Sentiment Analysis of Lyme Disease in Scientific Literature
Teo Susnjak. 2023 · 2023
Closest in time.
Is ChatGPT a Good Sentiment Analyzer? A Preliminary Study
Zengzhi Wang, Qiming Xie, Zixiang Ding, Yi Feng, and Rui Xia. 2023a · 2023
Closest in time.
Extractive Summarization via ChatGPT for Faithful Summary Generation
Haopeng Zhang, Xiao Liu, and Jiawei Zhang. 2023 · 2023
Closest in time.