Fetching the paper…
Reading the bibliography…
Dense document embeddings are central to neural retrieval.
Relevance feedback in information retrieval
J. J. Rocchio · 1971
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert, 2020
Omar Khattab and Matei Zaharia · 2004
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval, 2020
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk · 2007
Earlier work this paper cites.
The Probabilistic Relevance Framework: BM25 and Beyond
Stephen Robertson and Hugo Zaragoza · 2009
Earlier work this paper cites.
Simple English Wikipedia: A new text simplification task
William Coster and David Kauchak · 2011
Earlier work this paper cites.
Overcoming the lack of parallel data in sentence compression
Katja Filippova and Yasemin Altun · 2013
Earlier work this paper cites.
Open Question Answering Over Curated and Extracted Knowledge Bases
Anthony Fader, Luke Zettlemoyer, and Oren Etzioni · 2014
Earlier work this paper cites.
Identifying causal relations using parallel Wikipedia articles
Christopher Hidey and Kathy McKeown · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
news-please: A generic news crawler and extractor
Felix Hamborg, Norman Meuschke, Corinna Breitinger, and Bela Gipp · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset, 2018
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang · 2018
Earlier work this paper cites.
Conditional neural processes, 2018
Marta Garnelo, Dan Rosenbaum, Chris J. Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo J. Rezende, and S. M. Ali Eslami · 2018
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset, 2018
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
NPRF: A neural pseudo relevance feedback framework for ad-hoc information retrieval
Canjia Li, Yingfei Sun, Ben He, Le Wang, Kai Hui, Andrew Yates, Le Sun, and Jungang Xu · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Earlier work this paper cites.
Amazonqa: A review-based question answering task, 2019
Mansi Gupta, Nitish Kulkarni, Raghuveer Chanda, Anirudha Rayasam, and Zachary C Lipton · 2019
Earlier work this paper cites.
CodeSearchNet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Attentive neural processes, 2019
Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Jianmo Ni, Jiacheng Li, and Julian McAuley · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations, 2020
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering, 2020
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih · 2020
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models, 2020
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Cited alongside, same era.
S2orc: The semantic scholar open research corpus, 2020
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Dan S. Weld · 2020
Cited alongside, same era.
Passage re-ranking with bert, 2020
Rodrigo Nogueira and Kyunghyun Cho · 2020
Cited alongside, same era.
Unsupervised corpus aware language model pre-training for dense passage retrieval, 2021
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Later among the works it cites.
Colbertv2: Effective and efficient retrieval via lightweight late interaction, 2022
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia · 2022
Later among the works it cites.
Laprador: Unsupervised pretrained dense retriever for zero-shot text retrieval, 2022
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley · 2022
Later among the works it cites.
Scaling expert language models with unsupervised domain discovery, 2023
Suchin Gururangan, Margaret Li, Mike Lewis, Weijia Shi, Tim Althoff, Noah A. Smith, and Luke Zettlemoyer · 2023
Later among the works it cites.
Test-time adaptation via self-training with nearest neighbor information, 2023
Minguk Jang, Sae-Young Chung, and Hye Won Chung · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luyu Gao and Jamie Callan · 2021
Cited alongside, same era.
Scaling deep contrastive learning batch size under memory limited setup, 2021
Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan · 2021
Cited alongside, same era.
Efficiently teaching an effective dense retriever with balanced topic aware sampling, 2021
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury · 2021
Cited alongside, same era.
Gooaq: Open question answering with diverse answer types, 2021
Daniel Khashabi, Amos Ng, Tushar Khot, Ashish Sabharwal, Hannaneh Hajishirzi, and Chris Callison-Burch · 2021
Cited alongside, same era.
Paq: 65 million probably-asked questions and what you can do with them, 2021
Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel · 2021
Cited alongside, same era.
Large dual encoders are generalizable retrievers, 2021
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang · 2021
Cited alongside, same era.
Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering, 2021
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang · 2021
Cited alongside, same era.
Later among the works it cites.
Towards general text embeddings with multi-stage contrastive learning, 2023
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang · 2023
Later among the works it cites.
Text embeddings reveal (almost) as much as text, 2023
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush · 2023
Later among the works it cites.
Transformer neural processes: Uncertainty-aware meta learning via sequence modeling, 2023
Tung Nguyen and Aditya Grover · 2023
Later among the works it cites.
Introducing embed v3, Nov 2023
Nils Reimers, Elliot Choi, Amr Kayid, Alekhya Nandula, Manoj Govindassamy, and Abdullah Elkady · 2023
Later among the works it cites.
Global selection of contrastive batches via optimization on sample permutations
Vin Sachidananda, Ziyi Yang, and Chenguang Zhu · 2023
Later among the works it cites.
One embedder, any task: Instruction-finetuned text embeddings, 2023
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu · 2023
Later among the works it cites.
Optimizing test-time query representations for dense retrieval, 2023
Mujeen Sung, Jungsoo Park, Jaewoo Kang, Danqi Chen, and Jinhyuk Lee · 2023
Later among the works it cites.
Simlm: Pre-training with representation bottleneck for dense passage retrieval, 2023
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2023
Later among the works it cites.
The faiss library, 2024
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Closest in time.
Wikimedia downloads, 2024
Wikimedia Foundation · 2024
Closest in time.
Mode: Clip data experts via clustering, 2024
Jiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li, Luke Zettlemoyer, Shih-Fu Chang, Wen-Tau Yih, and Hu Xu · 2024
Closest in time.
Nomic embed: Training a reproducible long context text embedder, 2024
Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar · 2024
Closest in time.
In-context pretraining: Language modeling beyond document boundaries, 2024
Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Gergely Szilvasy, Rich James, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Scott Yih, and Mike Lewis · 2024
Closest in time.
Gistembed: Guided in-sample selection of training negatives for text embedding fine-tuning, 2024
Aivin V. Solatorio · 2024
Closest in time.
Text embeddings by weakly-supervised contrastive pre-training, 2024
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2024
Closest in time.
C-pack: Packaged resources to advance general chinese embedding, 2024
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie · 2024
Closest in time.