Fetching the paper…
Reading the bibliography…
The work of neural retrieval so far focuses on ranking short texts and is challenged with long documents.
Simple local attentions remain competitive for long-context tasks
Wenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Scott Yih, and Yashar Mehdad. 2022 · 1986
Earlier work this paper cites.
Okapi at TREC-3
Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994 · 1994
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
TREC 2006 genomics track overview
William R. Hersh, Aaron M. Cohen, Phoebe M. Roberts, and Hari Krishna Rekapalli. 2006 · 2006
Earlier work this paper cites.
TREC 2007 genomics track overview
William R. Hersh, Aaron M. Cohen, Lynn Ruslen, and Phoebe M. Roberts. 2007 · 2007
Earlier work this paper cites.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
Gordon V. Cormack, Charles L. A. Clarke, and Stefan Büttcher. 2009 · 2009
Earlier work this paper cites.
CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
TopicRank: Graph-based topic ranking for keyphrase extraction
Adrien Bougouin, Florian Boudin, and Béatrice Daille. 2013 · 2013
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning thematic similarity metric from article sections using triplet networks
Liat Ein Dor, Yosi Mass, Alon Halfon, Elad Venezian, Ilya Shnayderman, Ranit Aharonov, and Noam Slonim. 2018 · 2018
Earlier work this paper cites.
Pytrec_eval: An extremely fast python interface to trec_eval
Christophe Van Gysel and Maarten de Rijke. 2018 · 2018
Earlier work this paper cites.
Higher-order coreference resolution with coarse-to-fine inference
Kenton Lee, Luheng He, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
ReQA: An evaluation for end-to-end answer retrieval models
Amin Ahmad, Noah Constant, Yinfei Yang, and Daniel Cer. 2019 · 2019
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2019 · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
BERT for coreference resolution: Baselines and analysis
Mandar Joshi, Omer Levy, Luke Zettlemoyer, and Daniel Weld. 2019 · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston. 2020 · 2020
Cited alongside, same era.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Later among the works it cites.
Shallow pooling for sparse labels
Negar Arabzadeh, Alexandra Vtyurina, Xinyi Yan, and Charles LA Clarke. 2022 · 2022
Later among the works it cites.
An analysis of fusion functions for hybrid retrieval
Sebastian Bruch, Siyu Gai, and Amir Ingber. 2022 · 2022
Later among the works it cites.
Sedr: Segment representation learning for long documents dense retrieval
Junying Chen, Qingcai Chen, Dongfang Li, and Yutao Huang. 2022 · 2022
Later among the works it cites.
From distillation to hard negative sampling: Making sparse neural IR models more effective
Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SpanBERT: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
A dataset of information-seeking questions and answers anchored in research papers
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
SPLADE v2: Sparse lexical and expansion model for information retrieval
Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2021 · 2021
Cited alongside, same era.
DeCLUTR: Deep contrastive learning for unsupervised textual representations
John Giorgi, Osvald Nitski, Bo Wang, and Gary Bader. 2021 · 2021
Cited alongside, same era.
Pyserini: A python toolkit for reproducible information retrieval research with sparse and dense representations
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Frassetto Nogueira. 2021 · 2021
Cited alongside, same era.
Dense hierarchical retrieval for open-domain question answering
Ye Liu, Kazuma Hashimoto, Yingbo Zhou, Semih Yavuz, Caiming Xiong, and Philip Yu. 2021 · 2021
Cited alongside, same era.
Unsupervised corpus aware language model pre-training for dense passage retrieval
Luyu Gao and Jamie Callan. 2022 · 2022
Later among the works it cites.
On survivorship bias in ms marco
Prashansa Gupta and Sean MacAvaney. 2022 · 2022
Later among the works it cites.
ColBERTv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Later among the works it cites.
ConditionalQA: A complex reading comprehension dataset with conditional answers
Haitian Sun, William Cohen, and Ruslan Salakhutdinov. 2022 · 2022
Later among the works it cites.
RetroMAE: Pre-training retrieval-oriented language models via masked auto-encoder
Shitao Xiao, Zheng Liu, Yingxia Shao, and Zhao Cao. 2022 · 2022
Later among the works it cites.
Making a miracl: Multilingual information retrieval across a continuum of languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2022 · 2022
Later among the works it cites.
Hybrid hierarchical retrieval for open-domain question answering
Manoj Ghuhan Arivazhagan, Lan Liu, Peng Qi, Xinchi Chen, William Yang Wang, and Zhiheng Huang. 2023 · 2023
Closest in time.
Jina embeddings 2: 8192-token general-purpose text embeddings for long documents
Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, Maximilian Werk, Nan Wang, and Han Xiao. 2023 · 2023
Closest in time.
How to train your dragon: Diverse augmentation towards generalizable dense retrieval
Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023 · 2023
Closest in time.
Contextual masked auto-encoder for dense passage retrieval
Xing Wu, Guangyuan Ma, Meng Lin, Zijia Lin, Zhongyuan Wang, and Songlin Hu. 2023 · 2023
Closest in time.