Fetching the paper…
Reading the bibliography…
Recent dense retrievers increasingly leverage the robust text understanding capabilities of Large Language Models (LLMs), encoding queries and documents into a shared embedding space for effective retrieval.
MS MARCO: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of EMNLP . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Quora question pairs
DataCanary, hilfialkaff, Lili Jiang, Meg Risdal, Nikhil Dandekar, and tomtung. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of ACL . 1601–1611
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Pytrec_eval: An extremely fast python interface to trec_eval. In Proceedings of SIGIR . 873–876
Christophe Van Gysel and Maarten de Rijke. 2018 · 2018
Earlier work this paper cites.
DuReader: A chinese machine reading comprehension dataset from real-world applications. In Proceedings of the Workshop on Machine Reading for Question Answering . 37–46
Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, Xuan Liu, Tian Wu, and Haifeng Wang. 2018 · 2018
Earlier work this paper cites.
FEVER: A large-scale dataset for fact extraction and verification. In Proceedings of NAACL-HLT . 809–819
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of EMNLP . 2369–2380
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
ELI5: Long form question answering. In Proceedings of ACL . 3558–3567
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Natural Questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Open-domain question answering. In Proceedings of ACL . 34–37
Danqi Chen and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training. In Proceedings of ICML . 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering. In Proceedings of EMNLP . 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of SIGIR . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Fine-grained fact verification with kernel graph attention network. In Proceedings of ACL . 7342–7351
Zhenghao Liu, Chenyan Xiong, Maosong Sun, and Zhiyuan Liu. 2020 · 2020
Earlier work this paper cites.
SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of EMNLP . 6894–6910
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Earlier work this paper cites.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of NeurIPS
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval. In Proceedings of ICLR
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021 · 2021
Earlier work this paper cites.
Mr. TyDi: A multi-lingual benchmark for dense retrieval. In Proceedings of the 1st Workshop on Multilingual Representation Learning . 127–137
Xinyu Zhang, Xueguang Ma, Peng Shi, and Jimmy Lin. 2021 · 2021
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with io-awareness. In Proceedings of NeurIPS . 16344–16359
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models. In Proceedings of ICLR
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Sgpt: Gpt sentence embeddings for semantic search
Niklas Muennighoff. 2022 · 2022
Cited alongside, same era.
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al · 2022
Cited alongside, same era.
Large dual encoders are generalizable retrievers. In Proceedings of EMNLP . 9844–9855
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022 · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
MS MARCO web search: A large-scale information-rich web dataset with millions of real click labels. In Proceedings of WWW . 292–301
Qi Chen, Xiubo Geng, Corby Rosset, Carolyn Buractaon, Jingwen Lu, Tao Shen, Kun Zhou, Chenyan Xiong, Yeyun Gong, Paul Bennett, et al · 2024
Later among the works it cites.
Navigate through enigmatic labyrinth a survey of chain of thought reasoning: Advances, frontiers and future. In Proceedings of ACL . 1173–1203
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu. 2024 · 2024
Later among the works it cites.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024 · 2024
Later among the works it cites.
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies. In Proceedings of COLM
Shengding Hu, Yuge Tu, Xu Han, Ganqu Cui, Chaoqun He, Weilin Zhao, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Xinrong Zhang, Zhen Leng Thai, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, dahai li, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
COCO-DR: Combating the distribution shift in zero-shot dense retrieval with contrastive and distributionally robust learning. In Proceedings of EMNLP . 1462–1479
Yue Yu, Chenyan Xiong, Si Sun, Chao Zhang, and Arnold Overwijk. 2022 · 2022
Cited alongside, same era.
Multi-view document representation learning for open-domain dense retrieval. In Proceedings of ACL . 5990–6000
Shunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang, and Nan Duan. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Task-aware retrieval with instructions. In Findings of ACL . 3650–3675
Akari Asai, Timo Schick, Patrick Lewis, Xilun Chen, Gautier Izacard, Sebastian Riedel, Hannaneh Hajishirzi, and Wen-tau Yih. 2023 · 2023
Cited alongside, same era.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023 · 2023
Cited alongside, same era.
Precise zero-shot dense retrieval without relevance labels. In Proceedings of ACL . 1762–1777
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023 · 2023
Cited alongside, same era.
One embedder, any task: Instruction-finetuned text embeddings. In Findings of ACL . 1102–1121
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Leveraging llms for unsupervised dense retriever ranking. In Proceedings of SIGIR . 1307–1317
Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, and Guido Zuccon. 2024 · 2024
Later among the works it cites.
Think-to-talk or talk-to-think? When llms come up with an answer in multi-step reasoning
Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Ana Brassard, Keisuke Sakaguchi, and Kentaro Inui. 2024 · 2024
Later among the works it cites.
Gecko: Versatile text embeddings distilled from large language models
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, et al · 2024
Later among the works it cites.
Llama2vec: Unsupervised adaptation of large language models for dense retrieval. In Proceedings of ACL . 3490–3500
Chaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao, and Defu Lian. 2024 · 2024
Later among the works it cites.
Large language models as foundations for next-gen dense retrieval: A comprehensive empirical assessment. In Proceedings of EMNLP . 1354–1365
Kun Luo, Minghao Qin, Zheng Liu, Shitao Xiao, Jun Zhao, and Kang Liu. 2024 · 2024
Later among the works it cites.
Fine-tuning llama for multi-stage text retrieval. In Proceedings of SIGIR . 2421–2425
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
LLMs are also effective embedding models: An in-depth overview
Chongyang Tao, Tao Shen, Shen Gao, Junshuo Zhang, Zhen Li, Zhengwei Tao, and Shuai Ma. 2024 · 2024
Later among the works it cites.
Improving text embeddings with large language models. In Proceedings of ACL . 11897–11916
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024a · 2024
Later among the works it cites.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024b · 2024
Later among the works it cites.
Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information. In Findings of EMNLP . 6644–6653
Yongheng Zhang, Qiguang Chen, Jingxuan Zhou, Peng Wang, Jiasheng Si, Jin Wang, Wenpeng Lu, and Libo Qin. 2024 · 2024
Later among the works it cites.
Dense text retrieval based on pretrained language models: A survey
Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024 · 2024
Later among the works it cites.
PromptReps: Prompting large language models to generate dense and sparse representations for zero-shot document retrieval. In Proceedings of EMNLP . 4375–4391
Shengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin, and Guido Zuccon. 2024 · 2024
Later among the works it cites.
SmolLM2: When Smol Goes Big – Data-Centric Training of a Small Language Model
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, et al · 2025
Closest in time.
Making text embedders few-shot learners. In Proceedings of ICLR
Chaofan Li, Minghao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Yingxia Shao, Defu Lian, and Zheng Liu. 2025 · 2025
Closest in time.
Repetition improves language model embeddings. In Proceedings of ICLR
Jacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig, and Aditi Raghunathan. 2025 · 2025
Closest in time.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, et al · 2025
Closest in time.