Fetching the paper…
Reading the bibliography…
Current state-of-the-art document retrieval solutions mainly follow an index-retrieve paradigm, where the index is hard to be directly optimized for the final retrieval target.
Multidimensional binary search trees used for associative searching
Jon Louis Bentley · 1975
Earlier work this paper cites.
Algorithm as 136: A k-means clustering algorithm
John A Hartigan and Manchek A Wong · 1979
Earlier work this paper cites.
On relevance weights with little relevance information
Stephen E Robertson and Steve Walker · 1997
Earlier work this paper cites.
Document language models, query models, and risk minimization for information retrieval
John Lafferty and Chengxiang Zhai · 2001
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions
Mayur Datar, Nicole Immorlica, Piotr Indyk, and Vahab S Mirrokni · 2004
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson and Hugo Zaragoza · 2009
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Learning deep structured semantic models for web search using clickthrough data
Po-Sen Huang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Acero, and Larry Heck · 2013
Earlier work this paper cites.
Learning semantic representations using convolutional neural networks for web search
Yelong Shen, Xiaodong He, Jianfeng Gao, Li Deng, and Grégoire Mesnil · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al · 2015
Earlier work this paper cites.
A deep relevance matching model for ad-hoc retrieval
Jiafeng Guo, Yixing Fan, Qingyao Ai, and W Bruce Croft · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A comparison between term-based and embedding-based methods for initial retrieval
Tonglei Guo, Jiafeng Guo, Yixing Fan, Yanyan Lan, Jun Xu, and Xueqi Cheng · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Yu A Malkov and Dmitry A Yashunin · 2018
Earlier work this paper cites.
From neural re-ranking to neural ranking: Learning a sparse representation for inverted indexing
Hamed Zamani, Mostafa Dehghani, W Bruce Croft, Erik Learned-Miller, and Jaap Kamps · 2018
Cited alongside, same era.
Pre-training tasks for embedding-based large-scale retrieval
Wei-Cheng Chang, X Yu Felix, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar · 2019
Cited alongside, same era.
Context-aware sentence/passage term importance estimation for first stage retrieval
Zhuyun Dai and Jamie Callan · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Diskann: Fast accurate billion-point nearest neighbor search on a single node
Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi · 2019
Cited alongside, same era.
Parade: Passage representation aggregation for document reranking
Canjia Li, Andrew Yates, Sean MacAvaney, Ben He, and Yingfei Sun · 2020
Later among the works it cites.
Twinbert: Distilling knowledge to twin-structured compressed bert models for large-scale retrieval
Wenhao Lu, Jian Jiao, and Ruofei Zhang · 2020
Later among the works it cites.
Generation-augmented retrieval for open-domain question answering
Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen · 2020
Later among the works it cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2020
Later among the works it cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Cited alongside, same era.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
From doc2query to doctttttquery
Rodrigo Nogueira, Jimmy Lin, and AI Epistemic · 2019
Cited alongside, same era.
Document expansion by query prediction
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Cited alongside, same era.
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk · 2020
Later among the works it cites.
Is retriever merely an approximator of reader?
Sohee Yang and Minjoon Seo · 2020
Later among the works it cites.
Spann: Highly-efficient billion-scale approximate nearest neighbor search
Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang · 2021
Later among the works it cites.
Highly parallel autoregressive entity linking with discriminative correction
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Later among the works it cites.
Multilingual autoregressive entity linking
Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel, and Fabio Petroni · 2021
Later among the works it cites.
Condenser: a pre-training architecture for dense retrieval
Luyu Gao and Jamie Callan · 2021
Later among the works it cites.
Is your language model ready for dense representation fine-tuning
Luyu Gao and Jamie Callan · 2021
Later among the works it cites.
Unsupervised corpus aware language model pre-training for dense passage retrieval
Luyu Gao and Jamie Callan · 2021
Later among the works it cites.
Coil: Revisit exact lexical match in information retrieval with contextualized inverted list
Luyu Gao, Zhuyun Dai, and Jamie Callan · 2021
Later among the works it cites.
Sparse, dense, and attentional representations for text retrieval
Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins · 2021
Later among the works it cites.
End-to-end training of neural retrievers for open-domain question answering
Devendra Singh Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L Hamilton, and Bryan Catanzaro · 2021
Later among the works it cites.
Adversarial retriever-ranker for dense text retrieval
Hang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv, Nan Duan, and Weizhu Chen · 2021
Later among the works it cites.
Deep query likelihood model for information retrieval
Shengyao Zhuang, Hang Li, and G. Zuccon · 2021
Later among the works it cites.
Autoregressive search engines: Generating substrings as document identifiers
Michele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Wen-tau Yih, Sebastian Riedel, and Fabio Petroni · 2022
Closest in time.
Transformer memory as a differentiable search index
Yi Tay, Vinh Q Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al · 2022
Closest in time.