Fetching the paper…
Reading the bibliography…
Many use cases require retrieving smaller portions of text, and dense vector-based retrieval systems often perform better with shorter text segments, as the semantics are less likely to be over-compressed in the embeddings.
Approaches to passage retrieval in full text information systems
Gerard Salton, James Allan, and Chris Buckley · 1993
Earlier work this paper cites.
Passage-level evidence in document retrieval
James P Callan · 1994
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
FEVER: A Large-Scale Dataset for Fact Extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Goemotions: A dataset of fine-grained emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi · 2020
Earlier work this paper cites.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab and Matei Zaharia · 2020
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Cited alongside, same era.
BEIR: A Heterogeneous Benchmark for Zero-Shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Cited alongside, same era.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Cited alongside, same era.
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, et al · 2023
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
Rohan Jha, Bo Wang, Michael Günther, Saba Sturua, Mohammad Kalim Akram, and Han Xiao · 2024
Closest in time.
5 Levels of Text Splitting
Greg Kamradt · 2024
Closest in time.
Landmark embedding: A chunking-free embedding method for retrieval augmented long-context large language models
Kun Luo, Zheng Liu, Shitao Xiao, Tong Zhou, Yubo Chen, Jun Zhao, and Kang Liu · 2024
Closest in time.
Nomic Embed: Training a Reproducible Long Context Text Embedder
Zach Nussbaum, John X Morris, Brandon Duderstadt, and Andriy Mulyar · 2024
Closest in time.
jina-embeddings-v3: Multilingual Embeddings with Task LoRA
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
From Fixed-Size to NLP Chunking - A Deep Dive into Text Chunking Techniques
Krystian Safjan · 2023
Cited alongside, same era.
Introducing Contextual Retrieval, 2024
Anthropic · 2024
Cited alongside, same era.
Dense X retrieval: What retrieval granularity should we use?
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, and Dong Yu · 2024
Cited alongside, same era.
Closest in time.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Length-induced embedding collapse in transformer-based models
Yuqi Zhou, Sunhao Dai, Zhanshuo Cao, Xiao Zhang, and Jun Xu · 2024
Closest in time.
LongEmbed: Extending Embedding Models for Long Context Retrieval
Dawei Zhu, Liang Wang, Nan Yang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li · 2024
Closest in time.