Fetching the paper…
Reading the bibliography…
In this work we present a systematic empirical study focused on the suitability of the state-of-the-art multilingual encoders for cross-lingual document and sentence retrieval tasks across a number of diverse language pairs.
Psychometrika pp. 1–10 (1966)
Schönemann, P.H.: A generalized solution of the orthogonal Procrustes problem · 1966
Earlier work this paper cites.
In: Proceedings of SIGIR, pp. 232–241 (1994)
Robertson, S.E., Walker, S.: Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval · 1994
Earlier work this paper cites.
In: Proceedings of SIGIR, pp. 275–281 (1998)
Ponte, J.M., Croft, W.B.: A language modeling approach to information retrieval · 1998
Earlier work this paper cites.
In: Workshop of the Cross-Language Evaluation Forum for European Languages, pp. 44–63 (2003)
Braschler, M.: CLEF 2003–Overview of results · 2003
Earlier work this paper cites.
ACM Transactions on Information Systems pp. 179–214 (2004)
Zhai, C., Lafferty, J.: A study of smoothing methods for language models applied to information retrieval · 2004
Earlier work this paper cites.
In: Proceedings of the 10th Machine Translation Summit (MT SUMMIT), pp. 79–86 (2005)
Koehn, P.: Europarl: A parallel corpus for statistical machine translation · 2005
Earlier work this paper cites.
DOI https://doi.org/10.6028/NIST.SP.500-261
Voorhees, E.: Overview of the trec 2004 robust retrieval track (2005) · 2005
Earlier work this paper cites.
Mikolov, T., Le, Q.V., Sutskever, I.: Exploiting similarities among languages for machine translation · 2013
Earlier work this paper cites.
In: Proceedings of ADCS, pp. 3:1–3:8 (2015)
Hoogeveen, D., Verspoor, K.M., Baldwin, T.: CQADupStack: A benchmark data set for community question-answering research · 2015
Earlier work this paper cites.
In: Proceedings of SIGIR, pp. 363–372 (2015)
Vulić, I., Moens, M.F.: Monolingual and cross-lingual information retrieval models based on (bilingual) word embeddings · 2015
Earlier work this paper cites.
In: Proceedings of NAACL, pp. 1279–1289 (2016)
Lei, T., Joshi, H., Barzilay, R., Jaakkola, T., Tymoshenko, K., Moschitti, A., Màrquez, L.: Semi-supervised question retrieval with gated convolutions · 2016
Earlier work this paper cites.
Workshop on Cognitive Computing at NIPS (2016)
Nguyen, T., Rosenberg, M., Song, X., Gao, J., Tiwary, S., Majumder, R., Deng, L.: MS MARCO: A human generated machine reading comprehension dataset · 2016
Earlier work this paper cites.
In: Proceedings of LREC, pp. 3530–3534 (2016)
Ziemski, M., Junczys-Dowmunt, M., Pouliquen, B.: The United Nations parallel corpus v1.0 · 2016
Earlier work this paper cites.
In: Proceedings of SemEval, pp. 1–14 (2017)
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., Specia, L.: SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation · 2017
Earlier work this paper cites.
In: Proceedings of EMNLP, pp. 670–680 (2017)
Conneau, A., Kiela, D., Schwenk, H., Barrault, L., Bordes, A.: Supervised learning of universal sentence representations from natural language inference data · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1702.08734 (2017)
Johnson, J., Douze, M., Jégou, H.: Billion-scale similarity search with gpus · 2017
Earlier work this paper cites.
In: Proceedings of ICLR (2017)
Smith, S.L., Turban, D.H., Hamblin, S., Hammerla, N.Y.: Offline bilingual word vectors, orthogonal transformations and the inverted softmax · 2017
Earlier work this paper cites.
In: Proceedings of NeurIPS, pp. 5998–6008 (2017)
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need · 2017
Earlier work this paper cites.
In: Proceedings of ACL, pp. 789–798 (2018)
Artetxe, M., Labaka, G., Agirre, E.: A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings · 2018
Earlier work this paper cites.
In: Proceedings of EMNLP, pp. 169–174 (2018)
Cer, D., Yang, Y., Kong, S.y., Hua, N., Limtiaco, N., St. John, R., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Strope, B., Kurzweil, R.: Universal sentence encoder for English · 2018
Earlier work this paper cites.
In: Proceedings of LREC, pp. 1699–1704 (2018)
Conneau, A., Kiela, D.: SentEval: An evaluation toolkit for universal sentence representations · 2018
Earlier work this paper cites.
In: Proceedings of WMT, pp. 165–176 (2018)
Guo, M., Shen, Q., Yang, Y., Ge, H., Cer, D., Hernandez Abrego, G., Stevens, K., Constant, N., Sung, Y.H., Strope, B., Kurzweil, R.: Effective parallel corpus mining using bilingual sentence embeddings · 2018
Earlier work this paper cites.
In: Proceedings of NAACL, pp. 1112–1122 (2018)
Williams, A., Nangia, N., Bowman, S.: A broad-coverage challenge corpus for sentence understanding through inference · 2018
Earlier work this paper cites.
In: Proceedings of 11th Workshop on Building and Using Comparable Corpora, pp. 39–42 (2018)
Zweigenbaum, P., Sharoff, S., Rapp, R.: Overview of the third bucc shared task: Spotting parallel sentences in comparable corpora · 2018
Earlier work this paper cites.
In: Proceedings of EMNLP, pp. 3490–3496 (2019)
Akkalyoncu Yilmaz, Z., Yang, W., Zhang, H., Lin, J.: Cross-domain modeling of sentence-level evidence for document retrieval · 2019
Earlier work this paper cites.
Transactions of the ACL pp. 597–610 (2019)
Artetxe, M., Schwenk, H.: Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond · 2019
Earlier work this paper cites.
In: Proceedings of ACL: Workshop on Representation Learning for NLP, pp. 250–259 (2019)
Chidambaram, M., Yang, Y., Cer, D., Yuan, S., Sung, Y., Strope, B., Kurzweil, R.: Learning cross-lingual sentence representations via a multi-task dual-encoder model · 2019
Earlier work this paper cites.
In: Proceedings of NeurIPS, pp. 7059–7069 (2019)
Conneau, A., Lample, G.: Cross-lingual language model pretraining · 2019
Earlier work this paper cites.
In: Proceedings of SIGIR, pp. 985–988 (2019)
Dai, Z., Callan, J.: Deeper text understanding for IR with contextual neural language modeling · 2019
Earlier work this paper cites.
In: Proceedings of ACL, pp. 2978–2988 (2019)
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., Salakhutdinov, R.: Transformer-XL: Attentive language models beyond a fixed-length context · 2019
Cited alongside, same era.
In: Proceedings of NAACL, pp. 4171–4186 (2019)
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding · 2019
Cited alongside, same era.
In: Proceedings of EMNLP, pp. 55–65 (2019)
Ethayarajh, K.: How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings · 2019
Cited alongside, same era.
In: Proceedings of ACL, pp. 710–721 (2019)
Glavaš, G., Litschko, R., Ruder, S., Vulić, I.: How to (properly) evaluate cross-lingual word embeddings: On strong baselines, comparative analyses, and some misconceptions · 2019
Cited alongside, same era.
In: Proceedings of SIGIR, pp. 1109–1112 (2019)
Litschko, R., Glavaš, G., Vulić, I., Dietz, L.: Evaluating resource-lean cross-lingual embedding models in unsupervised retrieval · 2019
Cited alongside, same era.
In: Proceedings of LREC, p. 26 (2020)
Jiang, Z., El-Jaroudi, A., Hartmann, W., Karakos, D., Zhao, L.: Cross-lingual information retrieval with BERT · 2020
Later among the works it cites.
In: Proceedings of ICLR (2020)
Karthikeyan, K., Wang, Z., Mayhew, S., Roth, D.: Cross-lingual ability of multilingual BERT: An empirical study · 2020
Later among the works it cites.
In: Proceedings of SIGIR, p. 39–48 (2020)
Khattab, O., Zaharia, M.: Colbert: Efficient and effective passage search via contextualized late interaction over bert · 2020
Later among the works it cites.
In: Proceedings of EMNLP, pp. 4483–4499 (2020)
Lauscher, A., Ravishankar, V., Vulić, I., Glavaš, G.: From zero to hero: On the limitations of zero-shot language transfer with multilingual transformers · 2020
Later among the works it cites.
arXiv preprint arXiv:2008.09093 (2020)
Li, C., Yates, A., MacAvaney, S., He, B., Sun, Y.: Parade: Passage representation aggregation for document reranking · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, Q., McCarthy, D., Vulić, I., Korhonen, A.: Investigating cross-lingual alignment methods for contextualized embeddings with token-level evaluation · 2019
Cited alongside, same era.
arXiv preprint arXiv:1907.11692 (2019)
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V.: RoBERTa: A robustly optimized BERT pretraining approach · 2019
Cited alongside, same era.
In: Proceedings of SIGIR, pp. 1101–1104 (2019)
MacAvaney, S., Yates, A., Cohan, A., Goharian, N.: Cedr: Contextualized embeddings for document ranking · 2019
Cited alongside, same era.
arXiv preprint arXiv:1910.14424 (2019)
Nogueira, R., Yang, W., Cho, K., Lin, J.: Multi-stage document ranking with BERT · 2019
Cited alongside, same era.
In: Proceedings of ACL, pp. 4996–5001 (2019)
Pires, T., Schlinger, E., Garrette, D.: How multilingual is multilingual BERT? · 2019
Cited alongside, same era.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners (2019)
2019
Cited alongside, same era.
In: Proceedings of EMNLP, pp. 3973–3983 (2019)
Reimers, N., Gurevych, I.: Sentence-BERT: Sentence embeddings using siamese BERT-networks · 2019
Cited alongside, same era.
In: Proceedings of EMNLP, pp. 6008–6018 (2020)
Liang, Y., Duan, N., Gong, Y., Wu, N., Guo, F., Qi, W., Gong, M., Shou, L., Jiang, D., Cao, G., Fan, X., Zhang, R., Agrawal, R., Cui, E., Wei, S., Bharti, T., Qiao, Y., Chen, J.H., Wu, W., Liu, S., Yang, F., Campos, D., Majumder, R., Zhou, M.: XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation · 2020
Later among the works it cites.
arXiv preprint arXiv:2003.07278 (2020)
Liu, Q., Kusner, M.J., Blunsom, P.: A survey on contextual embeddings · 2020
Later among the works it cites.
In: Proceedings of EMNLP, pp. 4171–4179 (2020)
MacAvaney, S., Cohan, A., Goharian, N.: SLEDGE-Z: A zero-shot baseline for COVID-19 literature search · 2020
Later among the works it cites.
In: Proceedings of ECIR, pp. 246–254 (2020)
MacAvaney, S., Soldaini, L., Goharian, N.: Teaching a new dog old tricks: Resurrecting multilingual retrieval using zero-shot learning · 2020
Later among the works it cites.
In: Proceedings of EMNLP, pp. 2362–2376 (2020)
Ponti, E.M., Glavaš, G., Majewska, O., Liu, Q., Vulić, I., Korhonen, A.: XCOPA: A multilingual dataset for causal commonsense reasoning · 2020
Later among the works it cites.
In: Proceedings of EMNLP, pp. 4512–4525 (2020)
Reimers, N., Gurevych, I.: Making monolingual sentence embeddings multilingual using knowledge distillation · 2020
Later among the works it cites.
Transactions of the ACL pp. 842–866 (2020)
Rogers, A., Kovaleva, O., Rumshisky, A.: A primer in BERTology: What we know about how BERT works · 2020
Later among the works it cites.
In: Proceedings of EMNLP (Findings), pp. 2768–2773 (2020)
Shi, P., Bai, H., Lin, J.: Cross-lingual training of neural models for document ranking · 2020
Later among the works it cites.
In: Proceedings of EMNLP, pp. 7222–7240 (2020)
Vulić, I., Ponti, E.M., Litschko, R., Glavaš, G., Korhonen, A.: Probing pretrained language models for lexical semantics · 2020
Later among the works it cites.
In: Proceedings of ACL: System Demonstrations, pp. 87–94 (2020)
Yang, Y., Cer, D., Ahmad, A., Guo, M., Law, J., Constant, N., Abrego, G.H., Yuan, S., Tar, C., Sung, Y.h., Strope, B., Kurzweil, R.: Multilingual universal sentence encoder for semantic retrieval · 2020
Later among the works it cites.
In: Proceedings of SIGIR, p. 1637–1640 (2020)
Yu, P., Allan, J.: A study of neural matching models for cross-lingual IR · 2020
Later among the works it cites.
In: Proceedings of NeurIPS, pp. 17283–17297 (2020)
Zaheer, M., Guruganesh, G., Dubey, K.A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., Ahmed, A.: Big bird: Transformers for longer sequences · 2020
Later among the works it cites.
In: Proceedings of ACL, pp. 1656–1671 (2020)
Zhao, W., Glavaš, G., Peyrard, M., Gao, Y., West, R., Eger, S.: On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation · 2020
Later among the works it cites.
In: Proceedings of SIGIR, p. 1566–1576 (2021)
Craswell, N., Mitra, B., Yilmaz, E., Campos, D., Lin, J.: Ms marco: Benchmarking ranking models in the large-data regime · 2021
Closest in time.
Synthesis Lectures on Human Language Technologies. Morgan & Claypool Publishers (2021)
Lin, J., Nogueira, R., Yates, A.: Pretrained Transformers for Text Ranking: BERT and Beyond · 2021
Closest in time.
In: Proceedings of ECIR, pp. 342–358 (2021)
Litschko, R., Vulić, I., Ponzetto, S.P., Glavaš, G.: Evaluating multilingual text encoders for unsupervised cross-lingual retrieval · 2021
Closest in time.
In: Proceedings of EMNLP, pp. 1442–1459 (2021)
Liu, F., Vulić, I., Korhonen, A., Collier, N.: Fast, effective, and self-supervised: Transforming masked language models into universal lexical and sentence encoders · 2021
Closest in time.
In: Proceedings of NAACL, pp. 5835–5847 (2021)
Qu, Y., Ding, Y., Liu, J., Liu, K., Ren, R., Zhao, W.X., Dong, D., Wu, H., Wang, H.: RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering · 2021
Closest in time.
Shi, P., Zhang, R., Bai, H., Lin, J.: Cross-lingual training with dense retrieval for document retrieval · 2021
Closest in time.
In: Proceedings of NAACL, pp. 296–310 (2021)
Thakur, N., Reimers, N., Daxenberger, J., Gurevych, I.: Augmented SBERT: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks · 2021
Closest in time.
In: Proceedings of the 1st Workshop on Multilingual Representation Learning, pp. 127–137 (2021)
Zhang, X., Ma, X., Shi, P., Lin, J.: Mr. TyDi: A multi-lingual benchmark for dense retrieval · 2021
Closest in time.
In: Proceedings of ACL, pp. 5751–5767 (2021)
Zhao, M., Zhu, Y., Shareghi, E., Vulić, I., Reichart, R., Korhonen, A., Schütze, H.: A closer look at few-shot crosslingual transfer: The choice of shots matters · 2021
Closest in time.
In: Proceedings of *SEM, pp. 229–240 (2021)
Zhao, W., Eger, S., Bjerva, J., Augenstein, I.: Inducing language-agnostic multilingual representations · 2021
Closest in time.