Fetching the paper…
Reading the bibliography…
In this report, we introduce Piccolo2, an embedding model that surpasses other models in the comprehensive evaluation over 6 tasks on CMTEB benchmark, setting a new state-of-the-art.
A comparison and semi-quantitative analysis of words and character-bigrams as features in chinese text categorization
Jingyang Li, Maosong Sun, and Xian Zhang · 2006
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Onlineshoppingdataset, 2016
SophonPlus · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Multi-scale attentive interaction networks for chinese medical question answer selection
S. Zhang, X. Zhang, H. Wang, L. Guo, and S. Liu · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Ocnli: Original chinese natural language inference
Hai Hu, Kyle Richardson, Liang Xu, Lu Li, Sandra Kübler, and Lawrence S Moss · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Cord-19: The covid-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Douglas Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Kinney, et al · 2020
Earlier work this paper cites.
Clue: A chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, et al · 2020
Earlier work this paper cites.
mmarco: A multilingual version of the ms marco passage ranking dataset
Luiz Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, Roberto Lotufo, and Rodrigo Nogueira · 2021
Cited alongside, same era.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al · 2022
Cited alongside, same era.
Csl: A large-scale chinese scientific literature dataset
Yudong Li, Yuqing Zhang, Zhe Zhao, Linlin Shen, Weijie Liu, Weiquan Mao, and Hui Zhang · 2022
Nlizh dataset, 2023
shibing624 · 2023
Later among the works it cites.
Improving text embeddings with large language models
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighof · 2023
Later among the works it cites.
T2ranking: A large-scale chinese benchmark for passage ranking
Xiaohui Xie, Qian Dong, Bingning Wang, Feiyang Lv, Ting Yao, Weinan Gan, Zhijing Wu, Xiangsheng Li, Haitao Li, Yiqun Liu, et al · 2023
Later among the works it cites.
gte-qwen1.5-7b-instruct
alinlp · 2024
Closest in time.
aliyun-text-embedding
aliyun · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-cpr: A multi domain chinese dataset for passage retrieval
Dingkun Long, Qiong Gao, Kuan Zou, Guangwei Xu, Pengjun Xie, Ruijie Guo, Jian Xu, Guanjun Jiang, Luxi Xing, and Ping Yang · 2022
Cited alongside, same era.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Cited alongside, same era.
Dureader_retrieval: A large-scale chinese benchmark for passage retrieval from web search engine
Yifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen, Qiaoqiao She, Jing Liu, Hua Wu, and Haifeng Wang · 2022
Cited alongside, same era.
Cosent (1): A more effective sentence vector scheme than sentence bert, Jan 2022
Jianlin Su · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2022
Cited alongside, same era.
piccolo-large-zh
junqin huang · 2023
Cited alongside, same era.
Jdreview dataset, 2023
kuroneko5943 · 2023
Cited alongside, same era.
Closest in time.
baichuan-embedding, 2024
baichuan · 2024
Closest in time.
stella-mrl-large-zh-v3.5-1792d
infgrad · 2024
Closest in time.
acge-text-embedding
intsig · 2024
Closest in time.
Gecko: Versatile text embeddings distilled from large language models
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, et al · 2024
Closest in time.
text-embedding-v3
openai · 2024
Closest in time.
Sfr-embedding-mistral:enhance text retrieval with transfer learning
Ye Liu Rui Meng · 2024
Closest in time.
Open source strikes bread - new fluffy embeddings model, 2024
Darius Koenig Sean Lee, Aamir Shakir · 2024
Closest in time.