Fetching the paper…
Reading the bibliography…
As language-specific training data tends to be sparsely available compared to English, document retrieval in many languages has been largely relying on multilingual models.
Toward reproducible baselines: The open-source IR reproducibility challenge. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20–23, 2016. Proceedings 38 . Springer, 408–420
Jimmy Lin, Matt Crane, Andrew Trotman, Jamie Callan, Ishan Chattopadhyaya, John Foley, Grant Ingersoll, Craig Macdonald, and Sebastiano Vigna. 2016 · 2016
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Improving efficient neural ranking models with cross-architecture knowledge distillation
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2020 · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Teaching a new dog old tricks: Resurrecting multilingual retrieval using zero-shot learning. In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part II 42 . Springer, 246–254
Sean MacAvaney, Luca Soldaini, and Nazli Goharian. 2020 · 2020
Earlier work this paper cites.
LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 6442–6454
Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, and Yuji Matsumoto. 2020 · 2020
Earlier work this paper cites.
MMARCO: A multilingual version of the MS MMARCO passage ranking dataset
Luiz Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, Roberto Lotufo, and Rodrigo Nogueira. 2021 · 2021
Earlier work this paper cites.
BEIR: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
Pretrained transformers for text ranking: BERT and beyond. In Proceedings of the 14th ACM International Conference on web search and data mining . 1154–1156
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Earlier work this paper cites.
Optimizing dense retrieval model training with hard negatives. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1503–1512
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021 · 2021
Cited alongside, same era.
Mr. TyDi: A Multi-lingual Benchmark for Dense Retrieval. In Proceedings of the 1st Workshop on Multilingual Representation Learning . 127–137
Xinyu Zhang, Xueguang Ma, Peng Shi, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
ライブコンペティション:「AI 王~クイズ AI 日本一決定戦~」
鈴木 潤, 松田 耕史, 鈴木 正敏, 加藤 拓真, 宮脇 峻平, and 西田 京介. 2021 · 2021
Cited alongside, same era.
JGLUE: Japanese general language understanding evaluation. In Proceedings of the Thirteenth Language Resources and Evaluation Conference . 2957–2966
Kentaro Kurihara, Daisuke Kawahara, and Tomohide Shibata. 2022 · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022b · 2022
Later among the works it cites.
JCSE: Contrastive Learning of Japanese Sentence Embeddings and Its Applications
Zihao Chen, Hisashi Handa, and Kimiaki Shirahama. 2023 · 2023
Closest in time.
AnglE-optimized Text Embeddings
Xianming Li and Jing Li. 2023 · 2023
Closest in time.
Japanese SimCSE Technical Report
Hayato Tsukagoshi, Ryohei Sasano, and Koichi Takeda. 2023 · 2023
Closest in time.
C-pack: Packaged resources to advance general chinese embedding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022 · 2022
Cited alongside, same era.
Transfer learning approaches for building cross-language dense retrieval models. In European Conference on Information Retrieval . Springer, 382–396
Suraj Nair, Eugene Yang, Dawn Lawrie, Kevin Duh, Paul McNamee, Kenton Murray, James Mayfield, and Douglas W Oard. 2022 · 2022
Cited alongside, same era.
Squeezing water from a stone: a bag of tricks for further improving cross-encoder effectiveness for reranking. In European Conference on Information Retrieval . Springer, 655–670
Ronak Pradeep, Yuqi Liu, Xinyu Zhang, Yilin Li, Andrew Yates, and Jimmy Lin. 2022 · 2022
Cited alongside, same era.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 3715–3734
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Cited alongside, same era.
Progress in machine translation
Haifeng Wang, Hua Wu, Zhongjun He, Liang Huang, and Kenneth Ward Church. 2022a · 2022
Cited alongside, same era.
JaCWIR: Japanese Casual Web IR - 日本語情報検索評価のための小規模でカジュアルなWebタイトルと概要のデータセット
Yuichi Tateno. 2024a
Cited in the paper.
JQaRA: Japanese Question Answering with Retrieval Augmentation - 検索拡張(RAG)評価のための日本語Q&Aデータセット
Yuichi Tateno. 2024b
Cited in the paper.
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighof. 2023 · 2023
Closest in time.
MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2023 · 2023
Closest in time.
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024 · 2024
Closest in time.
SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . 370–390
Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, Aiti Aw, and Nancy Chen. 2024a · 2024
Closest in time.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024b · 2024
Closest in time.