Fetching the paper…
Reading the bibliography…
Pretrained multilingual text encoders based on neural Transformer architectures, such as multilingual BERT (mBERT) and XLM, have achieved strong performance on a myriad of language understanding tasks.
Ponte, J.M., Croft, W.B.: A language modeling approach to information retrieval. In: Proceedings of SIGIR. pp. 275–281 (1998)
1998
Earlier work this paper cites.
Braschler, M.: CLEF 2003–Overview of results. In: Workshop of the Cross-Language Evaluation Forum for European Languages. pp. 44–63 (2003)
2003
Earlier work this paper cites.
Zhai, C., Lafferty, J.: A study of smoothing methods for language models applied to information retrieval. ACM Transactions on Information Systems (TOIS) 22
2004
Earlier work this paper cites.
Koehn, P.: Europarl: A parallel corpus for statistical machine translation. In: Proceedings of the 10th Machine Translation Summit (MT SUMMIT). pp. 79–86 (2005)
2005
Earlier work this paper cites.
Hoogeveen, D., Verspoor, K.M., Baldwin, T.: CQADupStack: A benchmark data set for community question-answering research. In: Proceedings of ADCS. pp. 3:1–3:8 (2015)
2015
Earlier work this paper cites.
Vulić, I., Moens, M.F.: Monolingual and cross-lingual information retrieval models based on (bilingual) word embeddings. In: Proceedings of SIGIR. pp. 363–372 (2015)
2015
Earlier work this paper cites.
Lei, T., Joshi, H., Barzilay, R., Jaakkola, T., Tymoshenko, K., Moschitti, A., Màrquez, L.: Semi-supervised question retrieval with gated convolutions. In: Proceedings of NAACL. pp. 1279–1289 (2016)
2016
Earlier work this paper cites.
Ziemski, M., Junczys-Dowmunt, M., Pouliquen, B.: The United Nations parallel corpus v1.0. In: Proceedings of LREC. pp. 3530–3534 (2016)
2016
Earlier work this paper cites.
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., Specia, L.: SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In: Proceedings of SemEval. pp. 1–14 (2017)
2017
Earlier work this paper cites.
Conneau, A., Kiela, D., Schwenk, H., Barrault, L., Bordes, A.: Supervised learning of universal sentence representations from natural language inference data. In: Proceedings of EMNLP. pp. 670–680 (2017)
2017
Earlier work this paper cites.
Smith, S.L., Turban, D.H., Hamblin, S., Hammerla, N.Y.: Offline bilingual word vectors, orthogonal transformations and the inverted softmax. In: Proceedings of ICLR (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Proceedings of NeurIPS. pp. 5998–6008 (2017)
2017
Earlier work this paper cites.
Artetxe, M., Labaka, G., Agirre, E.: A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. In: Proceedings of ACL. pp. 789–798 (2018)
2018
Earlier work this paper cites.
Cer, D., Yang, Y., Kong, S.y., Hua, N., Limtiaco, N., St. John, R., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Strope, B., Kurzweil, R.: Universal sentence encoder for English. In: Proceedings of EMNLP. pp. 169–174 (2018)
2018
Earlier work this paper cites.
Conneau, A., Kiela, D.: SentEval: An evaluation toolkit for universal sentence representations. In: Proceedings of LREC. pp. 1699–1704 (2018)
2018
Earlier work this paper cites.
Conneau, A., Rinott, R., Lample, G., Williams, A., Bowman, S., Schwenk, H., Stoyanov, V.: XNLI: Evaluating cross-lingual sentence representations. In: Proceedings of EMNLP. pp. 2475–2485 (2018)
2018
Earlier work this paper cites.
Guo, M., Shen, Q., Yang, Y., Ge, H., Cer, D., Hernandez Abrego, G., Stevens, K., Constant, N., Sung, Y.H., Strope, B., Kurzweil, R.: Effective parallel corpus mining using bilingual sentence embeddings. In: Proceedings of WMT. pp. 165–176 (2018)
2018
Earlier work this paper cites.
Williams, A., Nangia, N., Bowman, S.: A broad-coverage challenge corpus for sentence understanding through inference. In: Proceedings of NAACL. pp. 1112–1122 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Artetxe, M., Schwenk, H.: Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the ACL pp. 597–610 (2019)
2019
Earlier work this paper cites.
Chidambaram, M., Yang, Y., Cer, D., Yuan, S., Sung, Y., Strope, B., Kurzweil, R.: Learning cross-lingual sentence representations via a multi-task dual-encoder model. In: Proceedings of the ACL Workshop on Representation Learning for NLP. pp. 250–259 (2019)
2019
Cited alongside, same era.
Conneau, A., Lample, G.: Cross-lingual language model pretraining. In: Proceedings of NeurIPS, pp. 7059–7069 (2019)
2019
Cited alongside, same era.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL. pp. 4171–4186 (2019)
2019
Cited alongside, same era.
Ethayarajh, K.: How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In: Proceedings of EMNLP-IJCNLP. pp. 55–65 (2019)
2019
Cited alongside, same era.
2020
Later among the works it cites.
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. In: Proceedings of NeurIPS (2020)
2020
Later among the works it cites.
Cao, S., Kitaev, N., Klein, D.: Multilingual alignment of contextual word representations. In: Proceedings of ICLR (2020)
2020
Later among the works it cites.
Clark, K., Luong, M., Le, Q.V., Manning, C.D.: ELECTRA: Pre-training text encoders as discriminators rather than generators. In: Proceedings of ICLR (2020)
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glavaš, G., Litschko, R., Ruder, S., Vulić, I.: How to (properly) evaluate cross-lingual word embeddings: On strong baselines, comparative analyses, and some misconceptions. In: Proceedings of ACL. pp. 710–721 (2019)
2019
Cited alongside, same era.
Litschko, R., Glavaš, G., Vulić, I., Dietz, L.: Evaluating resource-lean cross-lingual embedding models in unsupervised retrieval. In: Proceedings of SIGIR. pp. 1109–1112 (2019)
2019
Cited alongside, same era.
Liu, Q., McCarthy, D., Vulić, I., Korhonen, A.: Investigating cross-lingual alignment methods for contextualized embeddings with token-level evaluation. In: Proceedings of CoNLL. pp. 33–43 (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
MacAvaney, S., Yates, A., Cohan, A., Goharian, N.: Cedr: Contextualized embeddings for document ranking. In: Proceedings of SIGIR. pp. 1101–1104 (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Pires, T., Schlinger, E., Garrette, D.: How multilingual is multilingual BERT? In: Proceedings of ACL. pp. 4996–5001 (2019)
2019
Cited alongside, same era.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners. OpenAI blog 1
2019
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
Jiang, Z., El-Jaroudi, A., Hartmann, W., Karakos, D., Zhao, L.: Cross-lingual information retrieval with BERT. In: Proceedings of LREC. p. 26 (2020)
2020
Later among the works it cites.
Karthikeyan, K., Wang, Z., Mayhew, S., Roth, D.: Cross-lingual ability of multilingual BERT: An empirical study. In: Proceedings of ICLR (2020)
2020
Later among the works it cites.
Liang, Y., Duan, N., Gong, Y., Wu, N., Guo, F., Qi, W., Gong, M., Shou, L., Jiang, D., Cao, G., et al.: XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation. In: Proceedings of EMNLP (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
MacAvaney, S., Soldaini, L., Goharian, N.: Teaching a new dog old tricks: Resurrecting multilingual retrieval using zero-shot learning. In: Proceedings of ECIR. pp. 246–254 (2020)
2020
Later among the works it cites.
Ponti, E.M., Glavaš, G., Majewska, O., Liu, Q., Vulić, I., Korhonen, A.: XCOPA: A multilingual dataset for causal commonsense reasoning. In: Proceedings of EMNLP (2020)
2020
Later among the works it cites.
Reimers, N., Gurevych, I.: Making monolingual sentence embeddings multilingual using knowledge distillation. In: Proceedings of EMNLP (2020)
2020
Later among the works it cites.
Rogers, A., Kovaleva, O., Rumshisky, A.: A primer in BERTology: What we know about how BERT works. Transactions of the ACL (2020)
2020
Later among the works it cites.
Yang, Y., Cer, D., Ahmad, A., Guo, M., Law, J., Constant, N., Abrego, G.H., Yuan, S., Tar, C., Sung, Y.h., Strope, B., Kurzweil, R.: Multilingual universal sentence encoder for semantic retrieval. In: Proceedings of ACL: System Demonstrations. pp. 87–94 (2020)
2020
Later among the works it cites.
Yu, P., Allan, J.: A study of neural matching models for cross-lingual IR. In: Proceedings of SIGIR. p. 1637–1640 (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Zhao, W., Glavaš, G., Peyrard, M., Gao, Y., West, R., Eger, S.: On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation. In: Proceedings of ACL. pp. 1656–1671 (2020)
2020
Later among the works it cites.