Fetching the paper…
Reading the bibliography…
This article describes an efficient end-to-end speech translation (E2E-ST) framework based on non-autoregressive (NAR) models.
F. W. Stentiford and M. G. Steer, “Machine translation of speech,” British Telecom technology journal , vol. 6, no. 2, pp. 116–122, 1988
1988
Earlier work this paper cites.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural computation , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
A. Waibel, A. N. Jain, A. E. McNair, H. Saito, A. G. Hauptmann, and J. Tebelskis, “JANUS: a speech-to-speech translation system using connectionist and symbolic processing strategies,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on . IEEE Computer Society, 1991, pp. 793–796
1991
Earlier work this paper cites.
H. Ney, “Speech translation: Coupling of recognition and translation,” in Proceedings of ICASSP . IEEE, 1999, pp. 517–520
1999
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of ACL , 2002, pp. 311–318
2002
Earlier work this paper cites.
L. Qian, H. Zhou, Y. Bao, M. Wang, L. Qiu, W. Zhang, Y. Yu, and L. Li, “Glancing Transformer for non-autoregressive neural machine translation,” in Proceedings of ACL , 2021, pp. 1993–2003
2003
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin, and E. Herbst, “Moses: Open source toolkit for statistical machine translation,” in Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions , 2007, pp. 177–180
2007
Earlier work this paper cites.
C. Fügen, “A system for simultaneous translation of lectures and speeches,” Ph.D. dissertation, Karlsruhe Institute of Technology, 2009
2009
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in Proceedings of ASRU . IEEE, 2011
2011
Earlier work this paper cites.
M. Post, G. Kumar, A. Lopez, D. Karakos, C. Callison-Burch, and S. Khudanpur, “Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus,” in Proceedings of IWSLT , 2013
2013
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Proceedings of Interspeech , 2015, pp. 3586–3589
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Proceedings of ICLR , 2015
2015
Earlier work this paper cites.
A. Bérard, O. Pietquin, C. Servan, and L. Besacier, “Listen and translate: A proof of concept for end-to-end speech-to-text translation,” in Proceedings of NeurIPS 2016 End-to-end Learning for Speech and Audio Processing Workshop , 2016
2016
Earlier work this paper cites.
Y. Kim and A. M. Rush, “Sequence-level knowledge distillation,” in Proceedings of EMNLP , 2016, pp. 1317–1327
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of ACL , 2016, pp. 1715–1725
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of CVPR , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” in Proceedings of Interspeech , 2017, pp. 2625–2629
2017
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” in Proceedings of NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in Proceedings of NAACL-HLT , 2019, pp. 2012–2017
2017
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Earlier work this paper cites.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,” in Proceedings of ICASSP . IEEE, 2018, pp. 6224–6228
2018
Earlier work this paper cites.
J. Gu, J. Bradbury, C. Xiong, V. O. Li, and R. Socher, “Non-autoregressive neural machine translation,” in Proceedings of ICLR , 2018
2018
Earlier work this paper cites.
J. Lee, E. Mansimov, and K. Cho, “Deterministic non-autoregressive neural sequence modeling by iterative refinement,” in Proceedings of EMNLP , 2018, pp. 1173–1182
2018
Earlier work this paper cites.
J. Libovickỳ and J. Helcl, “End-to-end non-autoregressive neural machine translation with connectionist temporal classification,” in Proceedings of EMNLP , 2018, pp. 3016–3021
2018
Earlier work this paper cites.
L. Kaiser, S. Bengio, A. Roy, A. Vaswani, N. Parmar, J. Uszkoreit, and N. Shazeer, “Fast decoding in sequence models using discrete latent variables,” in Proceedings of ICML , 2018, pp. 2390–2399
2018
Earlier work this paper cites.
A. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg et al. , “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Proceedings of ICML , 2018, pp. 3918–3926
2018
Earlier work this paper cites.
A. C. Kocabiyikoglu, L. Besacier, and O. Kraif, “Augmenting Librispeech with French translations: A multimodal corpus for direct speech translation evaluation,” in Proceedings of LREC , 2018
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of EMNLP: System Demonstrations , 2018, pp. 66–71
2018
Earlier work this paper cites.
M. Post, “A call for clarity in reporting BLEU scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers , 2018, pp. 186–191
2018
Earlier work this paper cites.
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, “Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,” in Proceedings of NAACL-HLT , 2019, pp. 58–68
2019
Earlier work this paper cites.
Y. Liu, H. Xiong, Z. He, J. Zhang, H. Wu, H. Wang, and C. Zong, “End-to-end speech translation with knowledge distillation,” in Proceedings of Interspeech , 2019, pp. 1128–1132
2019
Earlier work this paper cites.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari, S. Laurenzo, and Y. Wu, “Leveraging weakly supervised data to improve end-to-end speech-to-text translation,” in Proceedings of ICASSP . IEEE, 2019, pp. 7180–7184
2019
Earlier work this paper cites.
J. Pino, L. Puzon, J. Gu, X. Ma, A. D. McCarthy, and D. Gopinath, “Harnessing indirect training data for end-to-end automatic speech translation: Tricks of the trade,” in Proceedings of IWSLT , 2019
2019
Cited alongside, same era.
J. Guo, X. Tan, D. He, T. Qin, L. Xu, and T.-Y. Liu, “Non-autoregressive neural machine translation with enhanced decoder input,” in Proceedings of AAAI , 2019, pp. 3723–3730
2019
Cited alongside, same era.
Z. Li, Z. Lin, D. He, F. Tian, T. Qin, L. Wang, and T.-Y. Liu, “Hint-based training for non-autoregressive machine translation,” in Proceedings of EMNLP , 2019, pp. 5708–5713
2019
Cited alongside, same era.
Y. Wang, F. Tian, D. He, T. Qin, C. Zhai, and T.-Y. Liu, “Non-autoregressive machine translation with auxiliary regularization,” in Proceedings of AAAI , 2019, pp. 5377–5384
2019
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for speech recognition,” in Proceedings of Interspeech , 2020, pp. 5036–5040
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Guo, L. Xu, and E. Chen, “Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation,” in Proceedings of ACL , 2020, pp. 376–385
2020
Later among the works it cites.
X. Kong, Z. Zhang, and E. Hovy, “Incorporating a local translation mechanism into non-autoregressive translation,” in Proceedings of EMNLP , 2020, pp. 1067–1073
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Wei, M. Wang, H. Zhou, J. Lin, and X. Sun, “Imitation learning for non-autoregressive neural machine translation,” in Proceedings of ACL , 2019, pp. 1304–1312
2019
Cited alongside, same era.
M. Ghazvininejad, O. Levy, Y. Liu, and L. Zettlemoyer, “Mask-predict: Parallel decoding of conditional masked language models,” in Proceedings of EMNLP , 2019, pp. 6112–6121
2019
Cited alongside, same era.
Z. Sun, Z. Li, H. Wang, D. He, Z. Lin, and Z. Deng, “Fast structured decoding for sequence models,” in Proceedings of NeurIPS , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds., 2019
2019
Cited alongside, same era.
J. Gu, C. Wang, and J. Zhao, “Levenshtein Transformer,” in Proceedings of NeurIPS , 2019, pp. 11 181–11 191
2019
Cited alongside, same era.
M. Stern, W. Chan, J. Kiros, and J. Uszkoreit, “Insertion Transformer: Flexible sequence generation via insertion operations,” in Proceedings of ICML , 2019, pp. 5976–5985
2019
Cited alongside, same era.
2019
Cited alongside, same era.
X. Ma, C. Zhou, X. Li, G. Neubig, and E. Hovy, “FlowSeq: Non-autoregressive conditional sequence generation with generative flow,” in Proceedings of EMNLP , 2019, pp. 4282–4292
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in Proceedings of NeurIPS , 2019, pp. 3171–3180
2019
Cited alongside, same era.
L. Ding, L. Wang, D. Wu, D. Tao, and Z. Tu, “Context-aware cross-attention for non-autoregressive translation,” in Proceedings of COLING , 2020, pp. 4396–4402
2020
Later among the works it cites.
P. Xie, Z. Cui, X. Chen, X. Hu, J. Cui, and B. Wang, “Infusing sequential information into conditional masked translation model with self-review mechanism,” in Proceedings of COLING , 2020, pp. 15–25
2020
Later among the works it cites.
H. Inaguma, S. Kiyono, K. Duh, S. Karita, N. Yalta, T. Hayashi, and S. Watanabe, “ESPnet-ST: All-in-one speech translation toolkit,” in Proceedings of ACL: System Demonstrations , 2020, pp. 302–311
2020
Later among the works it cites.
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, “Fairseq S2T: Fast speech-to-text modeling with fairseq,” in Proceedings of AACL: System Demonstrations , 2020, pp. 33–39
2020
Later among the works it cites.
Z. Sun and Y. Yang, “An EM approach to non-autoregressive conditional sequence generation,” in Proceedings of ICML , 2020, pp. 9249–9258
2020
Later among the works it cites.
M. Gaido, M. A. D. Gangi, M. Negri, M. Cettolo, and M. Turchi, “Contextualized translation of automatically segmented speech,” in Proceedings of Interspeech , 2020, pp. 1471–1475
2020
Later among the works it cites.
N.-Q. Pham, T.-L. Ha, T.-N. Nguyen, T.-S. Nguyen, E. Salesky, S. Stüker, J. Niehues, and A. Waibel, “Relative positional encoding for speech recognition and direct translation,” in Proceedings of Interspeech , 2020, pp. 31–35
2020
Later among the works it cites.
T. Potapczyk and P. Przybysz, “SRPOL’s system for the IWSLT 2020 end-to-end speech translation task,” in Proceedings of IWSLT , 2020, pp. 89–94
2020
Later among the works it cites.
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, “Pyannote.audio: neural building blocks for speaker diarization,” in Proceedings of ICASSP . IEEE, 2020, pp. 7124–7128
2020
Later among the works it cites.
L. Bentivogli, M. Cettolo, M. Gaido, A. Karakanta, A. Martinelli, M. Negri, and M. Turchi, “Cascade versus direct speech translation: Do the differences still make a difference?” in Proceedings of ACL , 2021, pp. 2873–2887
2021
Closest in time.
H. Inaguma, T. Kawahara, and S. Watanabe, “Source and target bidirectional knowledge distillation for end-to-end speech translation,” in Proceedings of NAACL-HLT , 2021, pp. 1872–1881
2021
Closest in time.
C. Du, Z. Tu, and J. Jiang, “Order-agnostic cross entropy for non-autoregressive machine translation,” in Proceedings of ICML , 2021
2021
Closest in time.
J. Gu and X. Kong, “Fully non-autoregressive neural machine translation: Tricks of the trade,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 120–133
2021
Closest in time.
Y. Higuchi, H. Inaguma, S. Watanabe, T. Ogawa, and T. Kobayashi, “Improved Mask-CTC for non-autoregressive end-to-end ASR,” in Proceedings of ICASSP . IEEE, 2021, pp. 8363–8367
2021
Closest in time.
H. Inaguma, Y. Higuchi, K. Duh, T. Kawahara, and S. Watanabe, “Orthros: Non-autoregressive end-to-end speech translation with dual-decoder,” in Proceedings of ICASSP . IEEE, 2021, pp. 7503–7507
2021
Closest in time.
S.-P. Chuang, Y.-S. Chuang, C.-C. Chang, and H.-y. Lee, “Investigating the reordering capability in CTC-based non-autoregressive end-to-end speech translation,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 1068–1077
2021
Closest in time.
M. Gaido, M. Cettolo, M. Negri, and M. Turchi, “CTC-based compression for direct speech translation,” in Proceedings of EACL , 2021, pp. 690–696
2021
Closest in time.
Q. Dong, M. Wang, H. Zhou, S. Xu, B. Xu, and L. Li, “Consecutive decoding for speech-to-text translation,” in Proceedings of AAAI , 2021
2021
Closest in time.
X. Zeng, L. Li, and Q. Liu, “RealTranS: End-to-end simultaneous speech translation with convolutional weighted-shrinking Transformer,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 2461–2474
2021
Closest in time.
P. Guo, F. Boyer, X. Chang, T. Hayashi, Y. Higuchi, H. Inaguma, N. Kamo, C. Li, D. Garcia-Romero, J. Shi et al. , “Recent developments on ESPnet toolkit boosted by Conformer,” in Proceedings of ICASSP . IEEE, 2021, pp. 5874–5878
2021
Closest in time.
Y. Hao, S. He, W. Jiao, Z. Tu, M. Lyu, and X. Wang, “Multi-task learning with shared encoder for non-autoregressive machine translation,” in Proceedings of NAACL-HLT , 2021, pp. 3989–3996
2021
Closest in time.
Q. Wang, H. Yu, S. Kuang, and W. Luo, “Hybrid-regressive neural machine translation,” 2021. [Online]. Available: https://openreview.net/forum?id=jYVY_piet7m
2021
Closest in time.
C. Zhao, M. Wang, Q. Dong, R. Ye, and L. Li, “NeurST: Neural speech translation toolkit,” in Proceedings of ACL: System Demonstrations , 2021, pp. 55–62
2021
Closest in time.
Q. Dong, R. Ye, M. Wang, H. Zhou, S. Xu, B. Xu, and L. Li, ““Listen, Understand and Translate”: Triple supervision decouples end-to-end speech-to-text translation,” in Proceedings of AAAI , 2021
2021
Closest in time.
C. Xu, B. Hu, Y. Li, Y. Zhang, S. Huang, Q. Ju, T. Xiao, and J. Zhu, “Stacked acoustic-and-textual encoding: Integrating the pre-trained models into speech translation encoders,” in Proceedings of ACL , 2021, pp. 2619–2630
2021
Closest in time.
H. Inaguma, B. Yan, S. Dalmia, P. Guo, J. Shi, K. Duh, and S. Watanabe, “ESPnet-ST IWSLT 2021 offline speech translation system,” in Proceedings of IWSLT , 2021, pp. 100–109
2021
Closest in time.
2021
Closest in time.
J. Kasai, N. Pappas, H. Peng, J. Cross, and N. Smith, “Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation,” in Proceedings of ICLR , 2021
2021
Closest in time.