Fetching the paper…
Reading the bibliography…
Sampling proper negatives from a large document pool is vital to effectively train a dense retrieval model.
Document expansion by query prediction
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019b · 1904
Earlier work this paper cites.
Repbert: Contextualized text embeddings for first-stage retrieval
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2020 · 2006
Earlier work this paper cites.
Is retriever merely an approximator of reader?
Sohee Yang and Minjoon Seo. 2020 · 2010
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Anserini: Enabling the use of lucene for information retrieval research
Peilin Yang, Hui Fang, and Jimmy Lin. 2017 · 2017
Earlier work this paper cites.
Training deep models faster with robust, approximate importance sampling
Tyler B Johnson and Carlos Guestrin. 2018 · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and François Fleuret. 2018 · 2018
Earlier work this paper cites.
Google dataset search: Building a search engine for datasets in an open web ecosystem
Dan Brickley, Matthew Burgess, and Natasha Noy. 2019 · 2019
Earlier work this paper cites.
Deeper text understanding for ir with contextual neural language modeling
Zhuyun Dai and Jamie Callan. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
Ambigqa: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020 · 2020
Cited alongside, same era.
Ernie 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2021
Later among the works it cites.
Measuring social biases in grounded vision and language embeddings
Candace Ross, Boris Katz, and Andrei Barbu. 2021 · 2021
Later among the works it cites.
End-to-end training of neural retrievers for open-domain question answering
Devendra Singh Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L. Hamilton, and Bryan Catanzaro. 2021 · 2021
Later among the works it cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Later among the works it cites.
Optimizing dense retrieval model training with hard negatives
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Is your language model ready for dense representation fine-tuning?
Luyu Gao and Jamie Callan. 2021 · 2021
Cited alongside, same era.
Scaling deep contrastive learning batch size under memory limited setup
Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan. 2021b · 2021
Cited alongside, same era.
Efficiently teaching an effective dense retriever with balanced topic aware sampling
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021 · 2021
Cited alongside, same era.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Édouard Grave. 2021 · 2021
Cited alongside, same era.
Sparse, dense, and attentional representations for text retrieval
Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021 · 2021
Cited alongside, same era.
Adversarial retriever-ranker for dense text retrieval
Hang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Unsupervised corpus aware language model pre-training for dense passage retrieval
Luyu Gao and Jamie Callan. 2022 · 2022
Closest in time.
Sentence-aware contrastive learning for open-domain passage retrieval
Wu Hong, Zhuosheng Zhang, Jinyuan Wang, and Hai Zhao. 2022 · 2022
Closest in time.
Yuxiang Lu, Yiding Liu, Jiaxiang Liu, Yunsheng Shi, Zhengjie Huang, Shikun Feng Yu Sun, Hao Tian, Hua Wu, Shuaiqiang Wang, Dawei Yin, et al. 2022 · 2022
Closest in time.
Curriculum contrastive context denoising for few-shot conversational dense retrieval
Kelong Mao, Zhicheng Dou, and Hongjin Qian. 2022 · 2022
Closest in time.
Domain-matched pre-training tasks for dense retrieval
Barlas Oğuz, Kushal Lakhotia, Anchit Gupta, Patrick Lewis, Vladimir Karpukhin, Aleksandra Piktus, Xilun Chen, Sebastian Riedel, Wen-tau Yih, Sonal Gupta, et al. 2022 · 2022
Closest in time.
Dureader_retrieval: A large-scale chinese benchmark for passage retrieval from web search engine
Yifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen, Qiaoqiao She, Jing Liu, Hua Wu, and Haifeng Wang. 2022 · 2022
Closest in time.
Learning to retrieve passages without supervision
Ori Ram, Gal Shachaf, Omer Levy, Jonathan Berant, and Amir Globerson. 2022 · 2022
Closest in time.
Laprador: Unsupervised pretrained dense retriever for zero-shot text retrieval
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2022 · 2022
Closest in time.