Fetching the paper…
Reading the bibliography…
As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 1909
Earlier work this paper cites.
CAIL2019-SCM: A dataset of similar case matching in legal domain
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Tianyang Zhang, Xianpei Han, Zhen Hu, Heng Wang, and Jianfeng Xu · 1911
Earlier work this paper cites.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vulic · 2003
Earlier work this paper cites.
CORD-19: the covid-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Chris Wilhelm, Boya Xie, Douglas Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier · 2004
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian J. McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Open question answering over curated and extracted knowledge bases
Anthony Fader, Luke Zettlemoyer, and Oren Etzioni · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
LCSTS: A large scale chinese short text summarization dataset
Baotian Hu, Qingcai Chen, and Fangze Zhu · 2015
Earlier work this paper cites.
A full-text learning to rank dataset for medical information retrieval
Vera Boteva, Demian Gholipour Ghalandari, Artem Sokolov, and Stefan Riezler · 2016
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Quora question pairs, 2017
DataCanary, hilfialkaff, Lili Jiang, Meg Risdal, Nikhil Dandekar, and tomtung · 2017
Earlier work this paper cites.
Searchqa: A new q&a dataset augmented with context from a search engine
Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Güney, Volkan Cirik, and Kyunghyun Cho · 2017
Earlier work this paper cites.
news-please - A generic news crawler and extractor
Felix Hamborg, Norman Meuschke, Corinna Breitinger, and Bela Gipp · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
The BQ corpus: A large-scale domain-specific chinese corpus for sentence semantic equivalence identification
Jing Chen, Qingcai Chen, Xin Liu, Haijun Yang, Daohe Lu, and Buzhou Tang · 2018
Earlier work this paper cites.
XNLI: evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
Dureader: a chinese machine reading comprehension dataset from real-world applications
Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, Xuan Liu, Tian Wu, and Haifeng Wang · 2018
Earlier work this paper cites.
LCQMC: A large-scale chinese question matching corpus
Xin Liu, Qingcai Chen, Chong Deng, Huajun Zeng, Jing Chen, Dongfang Li, and Buzhou Tang · 2018
Earlier work this paper cites.
Www’18 open challenge: Financial opinion mining and question answering
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
CARER: contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen · 2018
Earlier work this paper cites.
DRCD: a chinese machine reading comprehension dataset
Chih-Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Multi-scale attentive interaction networks for chinese medical question answer selection
Sheng Zhang, Xin Zhang, Hui Wang, Lixiang Guo, and Shanshan Liu · 2018
Earlier work this paper cites.
A span-extraction dataset for chinese machine reading comprehension
Yiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu · 2019
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Jianmo Ni, Jiacheng Li, and Julian J. McAuley · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Long and diverse text generation with planning-based hierarchical variational model
Zhihong Shao, Minlie Huang, Jiangtao Wen, Wenfei Xu, and Xiaoyan Zhu · 2019
Earlier work this paper cites.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge · 2019
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
OCNLI: original chinese natural language inference
Hai Hu, Kyle Richardson, Liang Xu, Lu Li, Sandra Kübler, and Lawrence S. Moss · 2020
Earlier work this paper cites.
Law question-answering dataset
Ustinian · 2020
Earlier work this paper cites.
TREC-COVID: constructing a pandemic information retrieval test collection
Ellen M. Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R. Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang · 2020
Cited alongside, same era.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Cited alongside, same era.
CLUE: A chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan · 2020
Cited alongside, same era.
mmarco: A multilingual version of MS MARCO passage ranking dataset
Luiz Henrique Bonifacio, Israel Campiotti, Roberto A. Lotufo, and Rodrigo Frassetto Nogueira · 2021
Cited alongside, same era.
Simcse: Simple contrastive learning of sentence embeddings
One embedder, any task: Instruction-finetuned text embeddings
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2023
Later among the works it cites.
T2ranking: A large-scale chinese benchmark for passage ranking
Xiaohui Xie, Qian Dong, Bingning Wang, Feiyang Lv, Ting Yao, Weinan Gan, Zhijing Wu, Xiangsheng Li, Haitao Li, Yiqun Liu, and Jin Ma · 2023
Later among the works it cites.
Ties-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel, and Mohit Bansal · 2023
Later among the works it cites.
Refgpt: Dialogue generation of gpt, by gpt, and for GPT
Dongjie Yang, Ruifeng Yuan, Yuantao Fan, Yifei Yang, Zili Wang, Shusen Wang, and Hai Zhao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Cited alongside, same era.
Xl-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Samin Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar · 2021
Cited alongside, same era.
Gooaq: Open question answering with diverse answer types
Daniel Khashabi, Amos Ng, Tushar Khot, Ashish Sabharwal, Hannaneh Hajishirzi, and Chris Callison-Burch · 2021
Cited alongside, same era.
Contractnli: A dataset for document-level natural language inference for contracts
Yuta Koreeda and Christopher D. Manning · 2021
Cited alongside, same era.
PAQ: 65 million probably-asked questions and what you can do with them
Patrick S. H. Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel · 2021
Cited alongside, same era.
MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark
Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2021
Cited alongside, same era.
I wish I would have loved this one, but I didn’t - A multilingual dataset for counterfactual detection in product review
James O’Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, and Danushka Bollegala · 2021
Cited alongside, same era.
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin · 2023
Later among the works it cites.
LIMA: less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Still need chunking when long-context models can do it all?, 2024
Jina AI · 2024
Later among the works it cites.
Scaling synthetic data creation with 1,000,000,000 personas
Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu · 2024
Later among the works it cites.
Dense X retrieval: What retrieval granularity should we use?
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, and Dong Yu · 2024
Later among the works it cites.
Mteb-french: Resources for french sentence embedding evaluation and analysis
Mathieu Ciancone, Imene Kerboua, Marion Schaeffer, and Wissam Siblini · 2024
Later among the works it cites.
Wikimedia downloads, 2024
Wikimedia Foundation · 2024
Later among the works it cites.
Bridging language and items for retrieval and recommendation
Yupeng Hou, Jiacheng Li, Zhankui He, An Yan, Xiusi Chen, and Julian J. McAuley · 2024
Later among the works it cites.
Piccolo2: General text embedding with multi-task hybrid loss training
Junqin Huang, Zhongjie Hu, Zihao Jing, Mengya Gao, and Yichao Wu · 2024
Later among the works it cites.
A survey on retrieval-augmented text generation for large language models
Yizheng Huang and Jimmy Huang · 2024
Later among the works it cites.
Nv-embed: Improved techniques for training llms as generalist embedding models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping · 2024
Later among the works it cites.
Landmark embedding: A chunking-free embedding method for retrieval augmented long-context large language models
Kun Luo, Zheng Liu, Shitao Xiao, Tong Zhou, Yubo Chen, Jun Zhao, and Kang Liu · 2024
Later among the works it cites.
Expertqa: Expert-curated questions and attributed answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth · 2024
Later among the works it cites.
Sfr-embedding-mistral:enhance text retrieval with transfer learning
Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz · 2024
Later among the works it cites.
Generative representational instruction tuning
Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela · 2024
Later among the works it cites.
PL-MTEB: polish massive text embedding benchmark
Rafal Poswiata, Slawomir Dadas, and Michal Perelkiewicz · 2024
Later among the works it cites.
Improving retrieval for RAG based question answering models on financial documents
Spurthi Setty, Katherine Jijo, Eden Chung, and Natan Vidra · 2024
Later among the works it cites.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O’Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzeminski, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Minh Vu Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker · 2024
Later among the works it cites.
jina-embeddings-v3: Multilingual embeddings with task lora
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, and Han Xiao · 2024
Later among the works it cites.
Thuctc: An efficient chinese text classifier, 2016
Maosong Sun, Jingyang Li, Zhipeng Guo, Yu Zhao, Yabin Zheng, Xiance Si, and Zhiyuan Liu · 2024
Later among the works it cites.
Large language models for data annotation and synthesis: A survey
Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu · 2024
Later among the works it cites.
Chinese semantic text similarity trainning dataset, 2016
Shancheng Tang, Yunyue Bai, and Fuyu Ma · 2024
Later among the works it cites.
Leveraging llms for synthesizing training data across many languages in multilingual dense retrieval
Nandan Thakur, Jianmo Ni, Gustavo Hernández Ábrego, John Wieting, Jimmy Lin, and Daniel Cer · 2024
Later among the works it cites.
Improving text embeddings with large language models
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei · 2024
Later among the works it cites.
C-pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie · 2024
Later among the works it cites.
Lm-cocktail: Resilient tuning of language models via model merging
Shitao Xiao, Zheng Liu, Peitian Zhang, and Xingrun Xing · 2024
Later among the works it cites.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan · 2024
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li · 2024
Later among the works it cites.
mgte: Generalized long-context text representation and reranking models for multilingual text retrieval
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang · 2024
Later among the works it cites.
Funnelrag: A coarse-to-fine progressive retrieval paradigm for RAG
Xinping Zhao, Yan Zhong, Zetian Sun, Xinshuo Hu, Zhenyu Liu, Dongfang Li, Baotian Hu, and Min Zhang · 2024
Later among the works it cites.
Opencodeinterpreter: Integrating code generation with execution and refinement
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue · 2024
Later among the works it cites.
Length-induced embedding collapse in transformer-based models
Yuqi Zhou, Sunhao Dai, Zhanshuo Cao, Xiao Zhang, and Jun Xu · 2024
Later among the works it cites.
Chatmed-dataset: An gpt generated medical query-response datasets for medcial large language models
Wei Zhu · 2024
Later among the works it cites.