Fetching the paper…
Reading the bibliography…
The black-box nature of end-to-end speech translation (E2E ST) systems makes it difficult to understand how source language inputs are being mapped to the target language.
“Interactive translation of conversational speech,”
A. Waibel, · 1996
Earlier work this paper cites.
“Moses: Open source toolkit for statistical machine translation,”
P. Koehn, H. Hoang, A. Birch, et al., · 2007
Earlier work this paper cites.
“Parallel implementations of word alignment tool,”
Q. Gao and S. Vogel, · 2008
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
R. Sennrich, B. Haddow, and A. Birch, · 2015
Earlier work this paper cites.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
A. Bérard, O. Pietquin, C. Servan, and L. Besacier, · 2016
Earlier work this paper cites.
“MuST-C: a Multilingual Speech Translation Corpus,”
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, · 2017
Earlier work this paper cites.
“Sequence-to-Sequence Models Can Directly Translate Foreign Speech,”
R. J. Weiss, J. Chorowski, N. Jaitly, et al., · 2017
Earlier work this paper cites.
“An operation sequence model for explainable neural machine translation,”
F. Stahlberg, D. Saunders, and B. Byrne, · 2018
Earlier work this paper cites.
“Espnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, et al., · 2018
Earlier work this paper cites.
“A call for clarity in reporting BLEU scores,”
M. Post, · 2018
Earlier work this paper cites.
“A comparative study on end-to-end speech to text translation,”
P. Bahar, T. Bieschke, and H. Ney, · 2019
Earlier work this paper cites.
“Attention-passing models for robust and data-efficient end-to-end speech translation,”
M. Sperber, G. Neubig, J. Niehues, and A. Waibel, · 2019
Earlier work this paper cites.
“Insertion-based decoding with automatically inferred generation order,”
J. Gu, Q. Liu, and K. Cho, · 2019
Earlier work this paper cites.
“Sequence modeling with unconstrained generation order,”
D. Emelianenko, E. Voita, and P. Serdyukov, · 2019
Earlier work this paper cites.
“The IWSLT 2019 KIT speech translation system,”
N.-Q. Pham, T.-S. Nguyen, T.-L. Ha, et al., · 2019
Earlier work this paper cites.
“STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework,”
M. Ma, L. Huang, H. Xiong, et al., · 2019
Cited alongside, same era.
“A comparative study on transformer vs RNN in speech applications,”
S. Karita, N. Chen, T. Hayashi, et al., · 2019
Cited alongside, same era.
“Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,”
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, · 2019
Cited alongside, same era.
“Synchronous speech recognition and speech-to-text translation with interactive decoding,”
Y. Liu, J. Zhang, H. Xiong, et al., · 2020
Cited alongside, same era.
“Dual-decoder transformer for joint automatic speech recognition and multilingual speech translation,”
H. Le, J. Pino, C. Wang, et al., · 2020
Cited alongside, same era.
“Searchable hidden intermediates for end-to-end models of decomposable sequence tasks,”
S. Dalmia, B. Yan, V. Raunak, et al., · 2021
Later among the works it cites.
“Consecutive decoding for speech-to-text translation,”
Q. Dong, M. Wang, H. Zhou, S. Xu, B. Xu, and L. Li, · 2021
Later among the works it cites.
“Streaming models for joint speech recognition and translation,”
O. Weller, M. Sperber, C. Gollan, and J. Kluivers, · 2021
Later among the works it cites.
“Guiding non-autoregressive neural machine translation decoding with reordering information,”
Q. Ran, Y. Lin, P. Li, and J. Zhou, · 2021
Later among the works it cites.
“Editor: an edit-based transformer with repositioning for neural machine translation with soft lexical constraints,”
W. Xu and M. Carpuat, · 2021
Later among the works it cites.
“Direct simultaneous speech-to-text translation assisted by synchronized streaming ASR,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Góis, K. Cho, and A. Martins, · 2020
Cited alongside, same era.
“Re-translation strategies for long form, simultaneous, spoken language translation,”
N. Arivazhagan, C. Cherry, I Te, W. Macherey, P. Baljekar, and G. Foster, · 2020
Cited alongside, same era.
“Consistent transcription and translation of speech,”
M. Sperber, H. Setiawan, et al., · 2020
Cited alongside, same era.
X. Ma, J. Pino, and P. Koehn, · 2020
Cited alongside, same era.
“ESPnet-ST: All-in-one speech translation toolkit,”
H. Inaguma, S. Kiyono, K. Duh, et al., · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
A. Gulati, J. Qin, C.-C. Chiu, et al., · 2020
Cited alongside, same era.
“ESPnet-ST IWSLT 2021 offline speech translation system,”
H. Inaguma, B. Yan, S. Dalmia, et al., · 2021
Cited alongside, same era.
J. Chen, M Ma, R. Zheng, and L. Huang, · 2021
Later among the works it cites.
“Findings of the IWSLT 2021 evaluation campaign,”
A. Anastasopoulos, O. Bojar, J. Bremerman, et al., · 2021
Later among the works it cites.
“Recent developments on espnet toolkit boosted by conformer,”
P. Guo, F. Boyer, X. Chang, et al., · 2021
Later among the works it cites.
“Streaming transformer asr with blockwise synchronous beam search,”
E. Tsunoo, Y. Kashiwagi, and S. Watanabe, · 2021
Later among the works it cites.
“The USTC-NELSLIP offline speech translation systems for IWSLT 2022,”
W. Zhang, Z. Ye, H. Tang, et al., · 2022
Closest in time.
“Revisiting end-to-end speech-to-text translation from scratch,”
B. Zhang, B. Haddow, and R. Sennrich, · 2022
Closest in time.
“CMU’s IWSLT 2022 dialect speech translation system,”
B. Yan, P. Fernandes, S. Dalmia, et al., · 2022
Closest in time.
“Ctc alignments improve autoregressive translation,”
B. Yan, S. Dalmia, Y. Higuchi, G. Neubig, F. Metze, A. Black, and S. Watanabe, · 2022
Closest in time.
“Learning when to translate for streaming speech,”
Q. Dong, Y. Zhu, M. Wang, and L. Li, · 2022
Closest in time.