Fetching the paper…
Reading the bibliography…
Current dense retrievers (DRs) are limited in their ability to effectively process misspelled queries, which constitute a significant portion of query traffic in commercial search engines.
Towards Robust Dense Retrieval via Local Ranking Alignment. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 . International Joint Conferences on Artificial Intelligence Organization, 1980–1986
Xuanang Chen, Jian Luo, Ben He, Le Sun, and Yingfei Sun. 2022a · 1986
Earlier work this paper cites.
The vocabulary problem in human-system communication
George W. Furnas, Thomas K. Landauer, Louis M. Gomez, and Susan T. Dumais. 1987 · 1987
Earlier work this paper cites.
“User revealment”—a comparison of initial queries and ensuing question development in online searching and in human reference interactions. In Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval . 11–18
Ragnar Nordlie. 1999 · 1999
Earlier work this paper cites.
Searching the web: The public and their queries
Amanda Spink, Dietmar Wolfram, Major BJ Jansen, and Tefko Saracevic. 2001 · 2001
Earlier work this paper cites.
Mining longitudinal Web queries: Trends and patterns
Peiling Wang, Michael W Berry, and Yiheng Yang. 2003 · 2003
Earlier work this paper cites.
Spelling correction in the PubMed search engine
W John Wilbur, Won Kim, and Natalie Xie. 2006 · 2006
Earlier work this paper cites.
Rank-biased precision for measurement of retrieval effectiveness
Alistair Moffat and Justin Zobel. 2008 · 2008
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
A large-scale query spelling correction corpus. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1261–1264
Matthias Hagen, Martin Potthast, Marcel Gohsen, Anja Rathgeber, and Benno Stein. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. In International Conference on Learning Representations
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters. In Proceedings of the 28th International Conference on Computational Linguistics . 6903–6915
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Earlier work this paper cites.
Improving efficient neural ranking models with cross-architecture knowledge distillation
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020 · 2020
Earlier work this paper cites.
Embedding-based retrieval in facebook search. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2553–2561
Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. 2020 · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Online, 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Distilling dense representations for ranking using tightly-coupled teachers
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Cited alongside, same era.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In International Conference on Learning Representations
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk. 2020 · 2020
Cited alongside, same era.
RepBERT: Contextualized text embeddings for first-stage retrieval
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2020 · 2020
Cited alongside, same era.
Condenser: a Pre-training Architecture for Dense Retrieval. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 981–993
Out-of-domain semantics to the rescue! zero-shot hybrid retrieval models. In Advances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022, Stavanger, Norway, April 10–14, 2022, Proceedings, Part I . Springer, 95–110
Tao Chen, Mingyang Zhang, Jing Lu, Michael Bendersky, and Marc Najork. 2022b · 2022
Later among the works it cites.
DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings. In Annual Conference of the North American Chapter of the Association for Computational Linguistics
Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljačić, Shang-Wen Li, Wen-Tau Yih, Yoon Kim, and James Glass. 2022 · 2022
Later among the works it cites.
Open Challenges in the Application of Dense Retrieval for Case Law Search
Pan Du, Hawre Hosseini, George Sanchez, and Filippo Pompili. 2022 · 2022
Later among the works it cites.
Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2843–2853
Luyu Gao and Jamie Callan. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luyu Gao and Jamie Callan. 2021 · 2021
Cited alongside, same era.
Rethink training of BERT rerankers in multi-stage retrieval pipeline. In European Conference on Information Retrieval . Springer, 280–286
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021 · 2021
Cited alongside, same era.
Efficiently teaching an effective dense retriever with balanced topic aware sampling. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 113–122
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021 · 2021
Cited alongside, same era.
In-batch negatives for knowledge distillation with tightly-coupled teachers for dense retrieval. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021) . 163–173
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
Less is More: Pretrain a Strong Siamese Encoder for Dense Text Retrieval Using a Weak Decoder. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 2780–2791
Shuqi Lu, Di He, Chenyan Xiong, Guolin Ke, Waleed Malik, Zhicheng Dou, Paul Bennett, Tie-Yan Liu, and Arnold Overwijk. 2021 · 2021
Cited alongside, same era.
How Deep is Your Learning: The DL-HARD Annotated Deep Learning Dataset
Iain Mackie, Jeffrey Dalton, and Andrew Yates. 2021 · 2021
Cited alongside, same era.
RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 5835–5847
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2021
Cited alongside, same era.
PAIR: Leveraging Passage-Centric Similarity Relation for Improving Dense Passage Retrieval. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 . 2173–2183
Ruiyang Ren, Shangwen Lv, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021a · 2021
Cited alongside, same era.
RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 2825–2835
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021b · 2021
Cited alongside, same era.
Tevatron: An Efficient and Flexible Toolkit for Dense Retrieval
Luyu Gao, Xueguang Ma, Jimmy J. Lin, and Jamie Callan. 2022 · 2022
Later among the works it cites.
Applications and Future of Dense Retrieval in Industry. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22) . Association for Computing Machinery, New York, NY, USA, 3373–3374
Yubin Kim. 2022 · 2022
Later among the works it cites.
RetroMAE: Pre-training Retrieval-oriented Transformers via Masked Auto-Encoder
Zheng Liu and Yingxia Shao. 2022 · 2022
Later among the works it cites.
Evaluating the robustness of retrieval pipelines with query variation generators. In European Conference on Information Retrieval . Springer, 397–412
Gustavo Penha, Arthur Câmara, and Claudia Hauff. 2022 · 2022
Later among the works it cites.
A thorough examination on zero-shot dense retrieval
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qifei Wu, Yuchen Ding, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2022 · 2022
Later among the works it cites.
LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval
Tao Shen, Xiubo Geng, Chongyang Tao, Can Xu, Xiaolong Huang, Binxing Jiao, Linjun Yang, and Daxin Jiang. 2022 · 2022
Later among the works it cites.
Analysing the Robustness of Dual Encoders for Dense Retrieval Against Misspellings (SIGIR ’22) . Association for Computing Machinery, New York, NY, USA, 2132–2136
Georgios Sidiropoulos and Evangelos Kanoulas. 2022 · 2022
Later among the works it cites.
Lecture Notes on Neural Information Retrieval
Nicola Tonellotto. 2022 · 2022
Later among the works it cites.
SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022 · 2022
Later among the works it cites.
Are neural ranking models robust?
Chen Wu, Ruqing Zhang, Jiafeng Guo, Yixing Fan, and Xueqi Cheng. 2022b · 2022
Later among the works it cites.
Contextual mask auto-encoder for dense passage retrieval
Xing Wu, Guangyuan Ma, Meng Lin, Zijia Lin, Zhongyuan Wang, and Songlin Hu. 2022a · 2022
Later among the works it cites.
Dense text retrieval based on pretrained language models: A survey
Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2022 · 2022
Later among the works it cites.
Robustness of Neural Rankers to Typos: A Comparative Study. In Proceedings of the 26th Australasian document computing symposium
Shengyao Zhuang, Xingyu Mao, and Guido Zuccon. 2022 · 2022
Later among the works it cites.
The tale of two MS MARCO–and their unfair comparisons
Carlos Lassance and Stéphane Clinchant. 2023 · 2023
Closest in time.