Fetching the paper…
Reading the bibliography…
How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It's well known that data augmentation is an efficient method to improve performance for many tasks by enlarging the dataset.
C. Dyer, V. Chahuneau, and N. A. Smith, “A simple, fast, and effective reparameterization of IBM model 2,” in Proc. of NAACL , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. of NeurIPS , 2017
2017
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi,” in Proc. of INTERSPEECH , 2017
2017
Earlier work this paper cites.
H. Zhang, M. Cissé, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in Proc. of ICLR , 2018
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proc. of EMNLP , 2018
2018
Earlier work this paper cites.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in Proc. of NAACL , 2019
2019
Earlier work this paper cites.
A. D. McCarthy, L. Puzon, and J. Pino, “Skinaugment: Auto-encoding speaker conversions for automatic speech translation,” in Proc. of ICASSP , 2020
2020
Cited alongside, same era.
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, “Fairseq S2T: Fast speech-to-text modeling with fairseq,” in Proc. of AACL , 2020
2020
Cited alongside, same era.
B. Zhang, P. Williams, I. Titov, and R. Sennrich, “Improving massively multilingual neural machine translation and zero-shot translation,” in Proc. of ACL , 2020
2020
Cited alongside, same era.
R. Ye, M. Wang, and L. Li, “End-to-end speech translation via cross-modal progressive training,” in Proc. of INTERSPEECH , 2021
2021
Cited alongside, same era.
C. Han, M. Wang, H. Ji, and L. Li, “Learning shared semantic space for speech-to-text translation,” in Proc. of ACL Findings , 2021
2021
S. Indurthi, M. A. Zaidi, N. K. Lakumarapu, B. Lee, H. Han, S. Ahn, S. Kim, C. Kim, and I. Hwang, “Task aware multi-task learning for speech to text tasks,” in Proc. of ICASSP , 2021
2021
Later among the works it cites.
H. Zhang, D. Qu, K. Shao, and X. Yang, “Dropdim: A regularization method for transformer networks,” IEEE Signal Processing Letters , 2022
2022
Closest in time.
T. K. Lam, S. Schamoni, and S. Riezler, “Sample, translate, recombine: Leveraging audio alignments for data augmentation in end-to-end speech translation,” in Proc. of ACL , 2022
2022
Closest in time.
2022
Closest in time.
Q. Fang, R. Ye, L. Li, Y. Feng, and M. Wang, “STEMM: Self-learning with speech-text manifold mixup for speech translation,” in Proc. of ACL , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
L. Meng, J. Xu, X. Tan, J. Wang, T. Qin, and B. Xu, “Mixspeech: Data augmentation for low-resource automatic speech recognition,” in Proc. of ICASSP , 2021
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
C. Xu, B. Hu, Y. Li, Y. Zhang, S. Huang, Q. Ju, T. Xiao, and J. Zhu, “Stacked acoustic-and-textual encoding: Integrating the pre-trained models into speech translation encoders,” in Proc. of ACL , 2021
2021
Cited alongside, same era.
P. Lison and J. Tiedemann, “OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles,” in Proc. LREC
Cited in the paper.
2022
Closest in time.
2022
Closest in time.
R. Ye, M. Wang, and L. Li, “Cross-modal contrastive learning for speech translation,” in Proc. of NAACL , 2022
2022
Closest in time.
Y. Du, Z. Zhang, W. Wang, B. Chen, J. Xie, and T. Xu, “Regularizing end-to-end speech translation with triangular decomposition agreement,” in Proc. of AAAI , 2022
2022
Closest in time.