Fetching the paper…
Reading the bibliography…
Recently, open-domain question answering systems have begun to rely heavily on annotated datasets to train neural passage retrievers.
Cumulated gain-based evaluation of ir techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen E. Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Pseudo-label : The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee. 2013 · 2013
Earlier work this paper cites.
Evaluation of baseline information retrieval for Polish open-domain question answering system
Michał Marcińczuk, Adam Radziszewski, Maciej Piasecki, Dominik Piasecki, and Marcin Ptak. 2013 · 2013
Earlier work this paper cites.
FastText.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Template-based question generation from retrieved sentences for improved unsupervised question answering
Alexander Fabbri, Patrick Ng, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020 · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
MLQA: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Cited alongside, same era.
KLEJ: Comprehensive benchmark for Polish language understanding
Piotr Rybak, Robert Mroczkowski, Janusz Tracz, and Ireneusz Gawlik. 2020 · 2020
Cited alongside, same era.
PAQ: 65 million probably-asked questions and what you can do with them
Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Later among the works it cites.
MKQA: A linguistically diverse benchmark for multilingual open domain question answering
Shayne Longpre, Yi Lu, and Joachim Daiber. 2021 · 2021
Later among the works it cites.
HerBERT: Efficiently pretrained transformer-based language model for Polish
Robert Mroczkowski, Piotr Rybak, Alina Wróblewska, and Ireneusz Gawlik. 2021 · 2021
Later among the works it cites.
RocketQAv2: A joint training method for dense passage retrieval and passage re-ranking
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, QiaoQiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021 · 2021
Later among the works it cites.
Tevatron: An efficient and flexible toolkit for dense retrieval
Luyu Gao, Xueguang Ma, Jimmy J. Lin, and Jamie Callan. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
Cited alongside, same era.
mMARCO: A multilingual version of MS MARCO passage ranking dataset
Luiz Henrique Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, , Roberto Lotufo, and Rodrigo Nogueira. 2021 · 2021
Cited alongside, same era.
MFAQ: a multilingual FAQ dataset
Maxime De Bruyn, Ehsan Lotfi, Jeska Buhmann, and Walter Daelemans. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, et al
Cited in the paper.
QA dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension
Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2022 · 2022
Later among the works it cites.
Improving question answering performance through manual annotation: Costs, benefits and strategies
Piotr Rybak, Piotr Przybyła, and Maciej Ogrodniczuk. 2022 · 2022
Later among the works it cites.