Fetching the paper…
Reading the bibliography…
Despite their recent popularity and well-known advantages, dense retrievers still lag behind sparse methods such as BM25 in their ability to reliably match salient phrases and rare entities in the query and to generalize to out-of-domain data.
Efficient inner product approximation in hybrid spaces
Xiang Wu, Ruiqi Guo, David Simcha, Dave Dopson, and Sanjiv Kumar. 2019 · 1903
Earlier work this paper cites.
Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval
S. E. Robertson and S. Walker. 1994 · 1994
Earlier work this paper cites.
A similarity measure for indefinite rankings
William Webber, Alistair Moffat, and Justin Zobel. 2010 · 2010
Earlier work this paper cites.
Apache lucene 4
Andrzej Bialecki, Robert Muir, and Grant Ingersoll. 2012 · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013 · 2013
Earlier work this paper cites.
Modeling of the question answering task in the yodaqa system
Petr Baudiš and Jan Šedivý. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
From doc2query to docTTTTTquery
Rodrigo Nogueira and Jimmy Lin. 2019 · 2019
Cited alongside, same era.
Context-aware term weighting for first stage passage retrieval
Zhuyun Dai and Jamie Callan. 2020 · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
Colbert: Efficient and effective passage search via contextualized late interaction over bert
Omar Khattab and Matei Zaharia. 2020 · 2020
Cited alongside, same era.
COIL: Revisit exact lexical match in information retrieval with contextualized inverted list
Learning passage impacts for inverted indexes
Antonio Mallia, Omar Khattab, Torsten Suel, and Nicola Tonellotto. 2021 · 2021
Closest in time.
Large dual encoders are generalizable retrievers
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2021 · 2021
Closest in time.
RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2021
Closest in time.
The curse of dense low-dimensional information retrieval for large index sizes
Nils Reimers and Iryna Gurevych. 2021 · 2021
Closest in time.
Simple entity-centric questions challenge dense retrievers
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021a · 2021
Cited alongside, same era.
Complement lexical retrieval model with semantic residual embeddings
Luyu Gao, Zhuyun Dai, Tongfei Chen, Zhen Fan, Benjamin Van Durme, and Jamie Callan. 2021b · 2021
Cited alongside, same era.
Efficiently teaching an effective dense retriever with balanced topic aware sampling
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021 · 2021
Cited alongside, same era.
Jimmy Lin and Xueguang Ma. 2021 · 2021
Cited alongside, same era.
In-batch negatives for knowledge distillation with tightly-coupled teachers for dense retrieval
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021b · 2021
Cited alongside, same era.
Sparse, dense, and attentional representations for text retrieval
Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021 · 2021
Cited alongside, same era.
Multi-task retrieval for knowledge-intensive tasks
Jean Maillard, Vladimir Karpukhin, Fabio Petroni, Wen-tau Yih, Barlas Oguz, Veselin Stoyanov, and Gargi Ghosh. 2021 · 2021
Cited alongside, same era.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Closest in time.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Closest in time.
xMoCo: Cross momentum contrastive learning for open-domain question answering
Nan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, and Linjun Yang. 2021 · 2021
Closest in time.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022 · 2022
Closest in time.
Another look at dpr: Reproduction of training and replication of retrieval
Xueguang Ma, Kai Sun, Ronak Pradeep, Minghan Li, and Jimmy Lin. 2022 · 2022
Closest in time.
Domain-matched pre-training tasks for dense retrieval
Barlas Oğuz, Kushal Lakhotia, Anchit Gupta, Patrick Lewis, Vladimir Karpukhin, Aleksandra Piktus, Xilun Chen, Sebastian Riedel, Wen tau Yih, Sonal Gupta, and Yashar Mehdad. 2022 · 2022
Closest in time.
Challenges in generalization in open domain question answering
Linqing Liu, Patrick Lewis, Sebastian Riedel, and Pontus Stenetorp. 2022 · 2029
Closest in time.