Fetching the paper…
Reading the bibliography…
End-to-end Speech-to-text Translation (E2E-ST), which directly translates source language speech to target language text, is widely useful in practice, but traditional cascaded approaches (ASR+MT) often suffer from error propagation in the pipeline.
Attention-Passing Models for Robust and Data-Efficient End-to-End Speech Translation
Sperber, M., Neubig, G., Niehues, J., and Waibel, A · 1904
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Speech translation: coupling of recognition and translation
Ney, H · 1999
Earlier work this paper cites.
On the integration of speech recognition and statistical machine translation
Matusov, E., Kanthak, S., and Ney, H · 2005
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
Statistical phrase-based speech translation
Mathias, L. and Byrne, W · 2006
Earlier work this paper cites.
A scalable method for preserving oral literature from small languages
Bird, S · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Povey, D., Ghoshal, A., Boulianne, G., Goel, N., Hannemann, M., Qian, Y., Schwarz, P., and Stemmer, G · 2011
Earlier work this paper cites.
Collecting bilingual audio in remote indigenous communities
Bird, S., Gawne, L., Gelbart, K., and McAlister, I · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
An unsupervised probability model for speech-to-translation alignment of low-resource languages
Anastasopoulos, A., Chiang, D., and Duong, L · 2016
Earlier work this paper cites.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Berard, A., Pietquin, O., Servan, C., and Besacier, L · 2016
Earlier work this paper cites.
Fma: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X · 2016
Earlier work this paper cites.
An attentional model for speech translation without transcription
Duong, L., Anastasopoulos, A., Chiang, D., Bird, S., and Cohn, T · 2016
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Cited alongside, same era.
Montreal forced aligner: Trainable text-speech alignment using kaldi
McAuliffe, M., Socolof, M., Mihuc, S., Wagner, M., and Sonderegger, M · 2017
Cited alongside, same era.
Hybrid ctc/attention architecture for end-to-end speech recognition
Watanabe, S., Hori, T., Kim, S., Hershey, J. R., and Hayashi, T · 2017
Cited alongside, same era.
Sequence-to-sequence models can directly translate foreign speech
Weiss, R. J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z · 2017
Cited alongside, same era.
Tied multitask learning for neural speech translation
MuST-C: a Multilingual Speech Translation Corpus
Di Gangi, M. A., Cattoni, R., Bentivogli, L., Negri, M., and Turchi, M · 2019
Later among the works it cites.
Adapting Transformer to End-to-End Spoken Language Translation
Gangi, M. A. D., Negri, M., and Turchi, M · 2019
Later among the works it cites.
End-to-End Speech Translation with Knowledge Distillation
Liu, Y., Xiong, H., Zhang, J., He, Z., Wu, H., Wang, H., and Zong, C · 2019
Later among the works it cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., and Le, Q. V · 2019
Later among the works it cites.
ERNIE: enhanced language representation with informative entities
Zhang, Z., Han, X., Liu, Z., Jiang, X., Sun, M., and Liu, Q · 2019
Later among the works it cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anastasopoulos, A. and Chiang, D · 2018
Cited alongside, same era.
End-to-end automatic speech translation of audiobooks
Berard, A., Besacier, L., Kocabiyikoglu, A., and Pietquin, O · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Cited alongside, same era.
End-to-end speech translation with the transformer
Vila, L. C., Escolano, C., Fonollosa, J. A. R., and Costa-jussà, M. R · 2018
Cited alongside, same era.
On using specaugment for end-to-end speech translation
Bahar, P., Zeyer, A., Schlüter, R., and Ney, H · 2019
Cited alongside, same era.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Bansal, S., Kamper, H., Livescu, K., Lopez, A., and Goldwater, S · 2019
Cited alongside, same era.
SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering
Chuang, Y.-S., Liu, C.-L., Lee, H.-Y., and Lee, L.-s · 2019
Cited alongside, same era.
Closest in time.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, H., Mohamed, A., and Auli, M · 2020
Closest in time.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Closest in time.
Espnet-st: All-in-one speech translation toolkit
Inaguma, H., Kiyono, S., Duh, K., Karita, S., Soplin, N. E. Y., Hayashi, T., and Watanabe, S · 2020
Closest in time.
Libri-light: A benchmark for asr with limited or no supervision
Kahn, J., Rivière, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P. E., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., Likhomanenko, T., Synnaeve, G., Joulin, A., Mohamed, A., and Dupoux, E · 2020
Closest in time.
Analyzing asr pretraining for low-resource speech-to-text translation
Stoian, M., Bansal, S., and Goldwater, S · 2020
Closest in time.
Bridging the gap between pre-training and fine-tuning for end-to-end speech translation
Wang, C., Wu, Y., Liu, S., Yang, Z., and Zhou, M · 2020
Closest in time.
Self-supervised representations improve end-to-end speech translation
Wu, A., Wang, C., Pino, J., and Gu, J · 2020
Closest in time.