Fetching the paper…
Reading the bibliography…
We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unlabeled speech and text in multiple languages.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Bojar, O., Buck, C., Callison-Burch, C., Federmann, C., Haddow, B., Koehn, P., Monz, C., Post, M., Soricut, R., and Specia, L · 2013
Earlier work this paper cites.
Speech recognition and keyword spotting for low-resource languages: Babel project research at cued
Gales, M. J. F., Knill, K., Ragni, A., and Rath, S. P · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Graves, A. and Jaitly, N · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Findings of the 2015 workshop on statistical machine translation
Bojar, O., Chatterjee, R., Federmann, C., Haddow, B., Huck, M., Hokamp, C., Koehn, P., Logacheva, V., Monz, C., Negri, M., Post, M., Scarton, C., Specia, L., and Turchi, M · 2015
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (WMT17)
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huang, S., Huck, M., Koehn, P., Liu, Q., Logacheva, V., Monz, C., Negri, M., Post, M., Rubino, R., Specia, L., and Turchi, M · 2017
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F., Wattenberg, M., Corrado, G., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Findings of the 2018 conference on machine translation (WMT18)
Bojar, O., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Koehn, P., and Monz, C · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Conneau, A., Rinott, R., Lample, G., Williams, A., Bowman, S. R., Schwenk, H., and Stoyanov, V · 2018
Earlier work this paper cites.
Zero-shot cross-lingual classification using multilingual neural machine translation, 2018
Eriguchi, A., Johnson, M., Firat, O., Kazawa, H., and Macherey, W · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
Lee, K., Firat, O., Agarwal, A., Fannjiang, C., and Sussillo, D · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2019
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G., Cherry, C., et al · 2019
Earlier work this paper cites.
Findings of the 2019 conference on machine translation (WMT19)
Barrault, L., Bojar, O., Costa-jussà, M. R., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., Malmasi, S., Monz, C., Müller, M., Pal, S., Post, M., and Zampieri, M · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Conneau, A. and Lample, G · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Large-scale multilingual speech recognition with a streaming end-to-end model
Kannan, A., Datta, A., Sainath, T. N., Weinstein, E., Ramabhadran, B., Wu, Y., Bapna, A., Chen, Z., and Lee, S · 2019
Cited alongside, same era.
Evaluating the cross-lingual effectiveness of massively multilingual neural machine translation
Siddhant, A., Johnson, M., Tsai, H., Ari, N., Riesa, J., Bapna, A., Firat, O., and Raman, K · 2020
Later among the works it cites.
Wang, Z., Tsvetkov, Y., Firat, O., and Cao, Y · 2020
Later among the works it cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Zhang, Y., Qin, J., Park, D. S., Han, W., Chiu, C.-C., Pang, R., Le, Q. V., and Wu, Y · 2020
Later among the works it cites.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Babu, A., Wang, C., Tjandra, A., Lakhotia, K., Xu, Q., Goyal, N., Singh, K., von Platen, P., Saraf, Y., Pino, J., et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Investigating multilingual nmt representations at scale
Kudugunta, S. R., Bapna, A., Caswell, I., Arivazhagan, N., and Firat, O · 2019
Cited alongside, same era.
Mlqa: Evaluating cross-lingual extractive question answering
Lewis, P., Oğuz, B., Rinott, R., Riedel, S., and Schwenk, H · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
Emerging cross-lingual structure in pretrained language models
Wu, S., Conneau, A., Li, H., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Common Voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, H., Mohamed, A., and Auli, M · 2020
Cited alongside, same era.
Findings of the 2020 conference on machine translation (WMT20)
Barrault, L., Biesialska, M., Bojar, O., Costa-jussà, M. R., Federmann, C., Graham, Y., Grundkiewicz, R., Haddow, B., Huck, M., Joanis, E., Kocmi, T., Koehn, P., Lo, C.-k., Ljubešić, N., Monz, C., Morishita, M., Nagata, M., Nakazawa, T., Pal, S., Post, M., and Zampieri, M · 2020
Cited alongside, same era.
Baevski, A., Hsu, W.-N., Conneau, A., and Auli, M · 2021
Later among the works it cites.
Joint unsupervised and supervised training for multilingual asr
Bai, J., Li, B., Zhang, Y., Bapna, A., Siddhartha, N., Sim, K. C., and Sainath, T. N · 2021
Later among the works it cites.
Slam: A unified encoder for speech and language modeling via speech-text joint pre-training
Bapna, A., Chung, Y.-a., Wu, N., Gulati, A., Jia, Y., Clark, J. H., Johnson, M., Riesa, J., Conneau, A., and Zhang, Y · 2021
Later among the works it cites.
w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., and Wu, Y · 2021
Later among the works it cites.
Multilingual and cross-lingual intent detection from spoken data
Gerz, D., Su, P.-H., Kusztos, R., Mondal, A., Lis, M., Singhal, E., Mrkšić, N., Wen, T.-H., and Vulić, I · 2021
Later among the works it cites.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Goyal, N., Gao, C., Chaudhary, V., Chen, P.-J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzman, F., and Fan, A · 2021
Later among the works it cites.
Explicit alignment objectives for multilingual bidirectional encoders
Hu, J., Johnson, M., Firat, O., Siddhant, A., and Neubig, G · 2021
Later among the works it cites.
nmt5 – is parallel data still relevant for pre-training massively multilingual language models?, 2021
Kale, M., Siddhant, A., Constant, N., Johnson, M., Al-Rfou, R., and Xue, L · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation, 2021
Li, J., Selvaraju, R. R., Gotmare, A. D., Joty, S., Xiong, C., and Hoi, S · 2021
Later among the works it cites.
Xtreme-r: Towards more challenging and nuanced multilingual evaluation
Ruder, S., Constant, N., Botha, J., Siddhant, A., Firat, O., Fu, J., Liu, P., Hu, J., Neubig, G., and Johnson, M · 2021
Later among the works it cites.
Self-training and pre-training are complementary for speech recognition
Xu, Q., Baevski, A., Likhomanenko, T., Tomasello, P., Conneau, A., Collobert, R., Synnaeve, G., and Auli, M · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Later among the works it cites.
Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation, 2021
Zheng, R., Chen, J., Ma, M., and Huang, L · 2021
Later among the works it cites.
When and why are pre-trained word embeddings useful for neural machine translation?
Qi, Y., Sachan, D., Felix, M., Padmanabhan, S., and Neubig, G · 2084
Closest in time.