Fetching the paper…
Reading the bibliography…
We present Mu$^{2}$SLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition (ASR), Automatic Speech Translation (AST) and Machine Translation (MT), in over 100 languages.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A · 2012
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Bojar, O., Buck, C., Callison-Burch, C., Federmann, C., Haddow, B., Koehn, P., Monz, C., Post, M., Soricut, R., and Specia, L · 2013
Earlier work this paper cites.
Speech recognition and keyword spotting for low-resource languages: Babel project research at cued
Gales, M. J. F., Knill, K., Ragni, A., and Rath, S. P · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Graves, A. and Jaitly, N · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Findings of the 2015 workshop on statistical machine translation
Bojar, O., Chatterjee, R., Federmann, C., Haddow, B., Huck, M., Hokamp, C., Koehn, P., Logacheva, V., Monz, C., Negri, M., Post, M., Scarton, C., Specia, L., and Turchi, M · 2015
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (WMT17)
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huang, S., Huck, M., Koehn, P., Liu, Q., Logacheva, V., Monz, C., Negri, M., Post, M., Rubino, R., Specia, L., and Turchi, M · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Weiss, R. J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z · 2017
Earlier work this paper cites.
Findings of the 2018 conference on machine translation (WMT18)
Bojar, O., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Koehn, P., and Monz, C · 2018
Earlier work this paper cites.
Devlin, J · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Post, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2019
Earlier work this paper cites.
Findings of the 2019 conference on machine translation (WMT19)
Barrault, L., Bojar, O., Costa-jussà, M. R., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., Malmasi, S., Monz, C., Müller, M., Pal, S., Post, M., and Zampieri, M · 2019
Earlier work this paper cites.
Robust neural machine translation with doubly adversarial inputs
Cheng, Y., Jiang, L., and Macherey, W · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Conneau, A. and Lample, G · 2019
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., and Le, Q. V · 2019
Cited alongside, same era.
Mass: Masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 2019
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., and Wu, Y · 2021
Later among the works it cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Later among the works it cites.
On generative spoken language modeling from raw audio
Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., et al · 2021
Later among the works it cites.
A general multi-task learning framework to leverage text data for speech to text tasks
Tang, Y., Pino, J., Wang, C., Ma, X., and Genzel, D · 2021
Later among the works it cites.
mslam: Massively multilingual joint pre-training for speech and text
Bapna, A., Cherry, C., Zhang, Y., Jia, Y., Johnson, M., Cheng, Y., Khanuja, S., Riesa, J., and Conneau, A · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Cited alongside, same era.
Findings of the 2020 conference on machine translation (WMT20)
Barrault, L., Biesialska, M., Bojar, O., Costa-jussà, M. R., Federmann, C., Graham, Y., Grundkiewicz, R., Haddow, B., Huck, M., Joanis, E., Kocmi, T., Koehn, P., Lo, C.-k., Ljubešić, N., Monz, C., Morishita, M., Nagata, M., Nakazawa, T., Pal, S., Post, M., and Zampieri, M · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning for speech recognition
Conneau, A., Baevski, A., Collobert, R., Mohamed, A., and Auli, M · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., et al · 2020
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Hu, J., Ruder, S., Siddhant, A., Neubig, G., Firat, O., and Johnson, M · 2020
Cited alongside, same era.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
Kasai, J., Pappas, N., Peng, H., Cross, J., and Smith, N. A · 2020
Cited alongside, same era.
Closest in time.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Closest in time.
Maestro-u: Leveraging joint speech-text representation learning for zero supervised speech asr
Chen, Z., Bapna, A., Rosenberg, A., Zhang, Y., Ramabhadran, B., Moreno, P., and Chen, N · 2022
Closest in time.
Multilingual mix: Example interpolation improves multilingual neural machine translation
Cheng, Y., Bapna, A., Firat, O., Cao, Y., Wang, P., and Macherey, W · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Closest in time.
Fleurs: Few-shot learning evaluation of universal representations of speech
Conneau, A., Ma, M., Khanuja, S., Zhang, Y., Axelrod, V., Dalmia, S., Riesa, J., Rivera, C., and Bapna, A · 2022
Closest in time.
Popuri, S., Chen, P.-J., Wang, C., Pino, J., Adi, Y., Gu, J., Hsu, W.-N., and Lee, A · 2022
Closest in time.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Closest in time.
Joist: A joint speech and text streaming model for asr
Sainath, T. N., Prabhavalkar, R., Bapna, A., Zhang, Y., Huo, Z., Chen, Z., Li, B., Wang, W., and Strohman, T · 2022
Closest in time.
Unified speech-text pre-training for speech translation and recognition
Tang, Y., Gong, H., Dong, N., Wang, C., Hsu, W.-N., Gu, J., Baevski, A., Li, X., Mohamed, A., Auli, M., et al · 2022
Closest in time.
Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition
Zhou, X., Wang, J., Cui, Z., Zhang, S., Yan, Z., Zhou, J., and Zhou, C · 2022
Closest in time.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Closest in time.
Google usm: Scaling automatic speech recognition beyond 100 languages
Zhang, Y., Han, W., Qin, J., Wang, Y., Bapna, A., Chen, Z., Chen, N., Li, B., Axelrod, V., Wang, G., et al · 2023
Closest in time.
When and why are pre-trained word embeddings useful for neural machine translation?
Qi, Y., Sachan, D., Felix, M., Padmanabhan, S., and Neubig, G · 2084
Closest in time.