Fetching the paper…
Reading the bibliography…
In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways.
R. C. Moore and W. Lewis, “Intelligent selection of language model training data,”
2010
Earlier work this paper cites.
K. Heafield, I. Pouzyrevsky, and et al., “Scalable modified Kneser-Ney language model estimation,” in
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to Sequence Learning with Neural Networks,” in
2014
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Earlier work this paper cites.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,” in
2018
Earlier work this paper cites.
A. Baevski and M. Auli, “Adaptive input representations for neural language modeling,”
2018
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised Pre-Training for Speech Recognition,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, “Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,” in
2019
Earlier work this paper cites.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari, S. Laurenzo, and Y. Wu, “Leveraging weakly supervised data to improve end-to-end speech-to-text translation,” in
2019
Earlier work this paper cites.
J. Pino, L. Puzon, J. Gu, X. Ma, A. D. McCarthy, and D. Gopinath, “Harnessing indirect training data for end-to-end automatic speech translation: Tricks of the trade,” in
2019
Earlier work this paper cites.
E. Salesky, M. Sperber, and A. W. Black, “Exploring phoneme-level speech representations for end-to-end speech translation,” in
2019
Earlier work this paper cites.
M. A. Di Gangi, M. Negri, and M. Turchi, “One-to-many multilingual end-to-end speech translation,” in
2019
Earlier work this paper cites.
H. Inaguma, K. Duh, T. Kawahara, and S. Watanabe, “Multilingual end-to-end speech translation,” in
2019
Cited alongside, same era.
S. B. H. K. K. Livescu and A. L. S. Goldwater, “Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,” in
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
A. Wu, C. Wang, J. Pino, and J. Gu, “Self-Supervised Representations Improve End-to-End Speech Translation,” in
2020
Later among the works it cites.
H. Nguyen, F. Bougares, N. Tomashenko, Y. Estève, and L. Besacier, “Investigating Self-Supervised Pre-Training for End-to-End Speech Translation,” in
2020
Later among the works it cites.
J. Pino, Q. Xu, X. Ma, M. J. Dousti, and Y. Tang, “Self-Training for End-to-End Speech Translation,” in
2020
Later among the works it cites.
C. Wang, J. Pino, A. Wu, and J. Gu, “Covost: A diverse multilingual speech-to-text translation corpus,” in
2020
Later among the works it cites.
C. Wang, A. Wu, and J. Pino, “Covost 2 and massively multilingual speech-to-text translation,”
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Kahn, A. Lee, and A. Hannun, “Self-training for end-to-end speech recognition,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Q. Xu, A. Baevski, T. Likhomanenko, P. Tomasello, A. Conneau, R. Collobert, G. Synnaeve, and M. Auli, “Self-training and pre-training are complementary for speech recognition,” in
2020
Cited alongside, same era.
2020
Later among the works it cites.
D. S. Park, Y. Zhang, Y. Jia, W. Han, C.-C. Chiu, and et al., “Improved noisy student training for automatic speech recognition,”
2020
Later among the works it cites.
J. Kahn, A. Lee, and A. Hannun, “Self-training for end-to-end speech recognition,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Iranzo-Sánchez, J. A. Silvestre-Cerdà, and et al., “Europarl-st: A multilingual corpus for speech translation of parliamentary debates,” in
2020
Later among the works it cites.
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” in
2020
Later among the works it cites.
G. Wenzek and et al., “CCNet: Extracting high quality monolingual datasets from web crawl data,” in
2020
Later among the works it cites.
B. Zoph, G. Ghiasi, T.-Y. Lin, Y. Cui, and et al., “Rethinking pre-training and self-training,”
2020
Later among the works it cites.
J. He, J. Gu, J. Shen, and M. Ranzato, “Revisiting self-training for neural sequence generation,” in
2020
Later among the works it cites.
Y. Tang, J. Pino, C. Wang, X. Ma, and D. Genzel, “A general multi-task learning framework to leverage text data for speech to text tasks,” in
2021
Closest in time.