Fetching the paper…
Reading the bibliography…
Speech translation models are unable to directly process long audios, like TED talks, which have to be split into shorter segments.
J. Sohn, N. S. Kim, and W. Sung, “A Statistical Model-Based Voice Activity Detection,” IEEE Signal Processing Letters , vol. 6, no. 1, pp. 1–3, 1999
1999
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: A Method for Automatic Evaluation of Machine Translation,” in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics , ser. ACL ’02. USA: Association for Computational Linguistics, 2002, p. 311–318
2002
Earlier work this paper cites.
E. Matusov, G. Leusch, O. Bender, and H. Ney, “Evaluating Machine Translation Output with Automatic Sentence Segmentation,” in Proceedings of the Second International Workshop on Spoken Language Translation , Pittsburgh, Pennsylvania, USA, Oct. 24-25 2005
2005
Earlier work this paper cites.
E. Matusov, D. Hillard, M. Magimai-Doss, D. Hakkani-Tur, M. Ostendorf, and H. Ney, “Improving speech translation with automatic boundary prediction,” 08 2007, pp. 2449–2452
2007
Earlier work this paper cites.
M. Sinclair, P. Bell, A. Birch, and F. McInnes, “A semi-Markov model for speech segmentation with an utterance-break prior,” in Proc. Interspeech 2014 , 2014, pp. 2351–2355
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in Proceedings of 3rd International Conference on Learning Representations, ICLR, San Diego, USA , 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 2012–2017
2017
Earlier work this paper cites.
E. Matusov, P. Wilken, P. Bahar, J. Schamper, P. Golik, A. Zeyer, J. A. Silvestre-Cerda, A. Martinez-Villaronga, H. Pesch, and J.-T. Peter, “Neural Speech Translation at AppTek,” in Proceedings of the 15th International Workshop on Spoken Language Translation , 2018, pp. 104–111
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Brussels, Belgium: Association for Computational Linguistics, Nov. 2018, pp. 66–71
2018
Cited alongside, same era.
M. Post, “A Call for Clarity in Reporting BLEU Scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers . Belgium, Brussels: Association for Computational Linguistics, Oct. 2018, pp. 186–191
2018
Cited alongside, same era.
E. Ansari, A. Axelrod, N. Bach, O. Bojar, R. Cattoni, F. Dalvi, N. Durrani, M. Federico, C. Federmann, J. Gu, F. Huang, K. Knight, X. Ma, A. Nagesh, M. Negri, J. Niehues, J. Pino, E. Salesky, X. Shi, S. Stüker, M. Turchi, A. Waibel, and C. Wang, “FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN,” in Proceedings of the 17th International Conference on Spoken Language Translation . Online: Association for Computational Linguistics, Jul. 2020, pp. 1–34
2020
Cited alongside, same era.
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, “fairseq S2T: Fast Speech-to-Text Modeling with fairseq,” in Proceedings of the 2020 Conference of the Asian Chapter of the Association for Computational Linguistics (AACL): System Demonstrations , 2020
2020
Later among the works it cites.
E. Salesky, M. Wiesner, J. Bremerman, R. Cattoni, M. Negri, M. Turchi, D. W. Oard, and M. Post, “The Multilingual TEDx Corpus for Speech Recognition and Translation,” in Proc. Interspeech 2021 , 2021, pp. 3655–3659
2021
Later among the works it cites.
L. Bentivogli, M. Cettolo, M. Gaido, A. Karakanta, A. Martinelli, M. Negri, and M. Turchi, “Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference?” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Online: Association for Computational Linguistics, Aug. 2021, pp. 2873–2887
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Potapczyk and P. Przybysz, “SRPOL’s System for the IWSLT 2020 End-to-End Speech Translation Task,” in Proceedings of the 17th International Conference on Spoken Language Translation . Online: Association for Computational Linguistics, Jul. 2020, pp. 89–94
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 12 449–12 460
2020
Cited alongside, same era.
J. Iranzo-Sánchez, J. A. Silvestre-Cerdà, J. Jorge, N. Roselló, A. Giménez, A. Sanchis, J. Civera, and A. Juan, “Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 8229–8233
2020
Cited alongside, same era.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-Art Natural Language Processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Online: Association for Computational Linguistics, Oct. 2020, pp. 38–45
2020
Cited alongside, same era.
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu, “On Layer Normalization in the Transformer Architecture,” in Proceedings of the 37th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 119. PMLR, 13–18 Jul 2020, pp. 10 524–10 533
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Anastasopoulos, O. Bojar, J. Bremerman, R. Cattoni, M. Elbayad, M. Federico, X. Ma, S. Nakamura, M. Negri, J. Niehues, J. Pino, E. Salesky, S. Stüker, K. Sudoh, M. Turchi, A. Waibel, C. Wang, and M. Wiesner, “FINDINGS OF THE IWSLT 2021 EVALUATION CAMPAIGN,” in Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT 2021) . Bangkok, Thailand (online): Association for Computational Linguistics, Aug. 2021, pp. 1–29
2021
Later among the works it cites.
M. Gaido, M. Negri, M. Cettolo, and M. Turchi, “Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation,” in Proceedings of The Fourth International Conference on Natural Language and Speech Processing (ICNLSP 2021) . Trento, Italy: Association for Computational Linguistics, 12–13 Nov. 2021, pp. 55–62
2021
Later among the works it cites.
2021
Later among the works it cites.
G. I. Gállego, I. Tsiamas, C. Escolano, J. A. R. Fonollosa, and M. R. Costa-jussà, “End-to-end Speech Translation with Pre-trained Models and Adapters: UPC at IWSLT 2021,” in Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT 2021) . Bangkok, Thailand (online): Association for Computational Linguistics, Aug. 2021, pp. 110–119
2021
Later among the works it cites.
Y. Tang, J. Pino, X. Li, C. Wang, and D. Genzel, “Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task,” in ACL , 2021
2021
Later among the works it cites.