Fetching the paper…
Reading the bibliography…
Fast inference speed is an important goal towards real-world deployment of speech translation (ST) systems.
“Speech translation: Coupling of recognition and translation,”
Hermann Ney, · 1999
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni et al., · 2002
Earlier work this paper cites.
“Moses: Open source toolkit for statistical machine translation,”
Philipp Koehn et al., · 2007
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey et al., · 2011
Earlier work this paper cites.
“Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus,”
Matt Post et al., · 2013
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko et al., · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma et al., · 2015
Earlier work this paper cites.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
Alexandre Bérard et al., · 2016
Earlier work this paper cites.
“Dynamic transcription for low-latency speech translation.,”
Jan Niehues et al., · 2016
Earlier work this paper cites.
“Sequence-level knowledge distillation,”
Yoon Kim et al., · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich et al., · 2016
Earlier work this paper cites.
“Unwritten languages demand attention too! word discovery with encoder-decoder models,”
Marcely Zanon Boito et al., · 2017
Earlier work this paper cites.
“Sequence-to-sequence models can directly translate foreign speech,”
Ron J Weiss et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani et al., · 2017
Earlier work this paper cites.
“MuST-C: a Multilingual Speech Translation Corpus,”
Mattia A. Di Gangi et al., · 2017
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe et al., · 2017
Earlier work this paper cites.
“End-to-end automatic speech translation of audiobooks,”
Alexandre Bérard et al., · 2018
Earlier work this paper cites.
“Non-autoregressive neural machine translation,”
Jiatao Gu et al., · 2018
Earlier work this paper cites.
“Deterministic non-autoregressive neural sequence modeling by iterative refinement,”
Jason Lee et al., · 2018
Earlier work this paper cites.
“Fast decoding in sequence models using discrete latent variables,”
Lukasz Kaiser et al., · 2018
Cited alongside, same era.
“End-to-end non-autoregressive neural machine translation with connectionist temporal classification,”
Jindřich Libovickỳ et al., · 2018
Cited alongside, same era.
“Parallel wavenet: Fast high-fidelity speech synthesis,”
Aaron Oord et al., · 2018
Cited alongside, same era.
“Augmenting Librispeech with French translations: A multimodal corpus for direct speech translation evaluation,”
Ali Can Kocabiyikoglu et al., · 2018
Cited alongside, same era.
“Subword regularization: Improving neural network translation models with multiple subword candidates,”
Taku Kudo, · 2018
Cited alongside, same era.
“Leveraging weakly supervised data to improve end-to-end speech-to-text translation,”
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park et al., · 2019
Later among the works it cites.
“BERT: Pre-training of deep bidirectional Transformers for language understanding,”
Jacob Devlin et al., · 2019
Later among the works it cites.
“Bridging the gap between pre-training and fine-tuning for end-to-end speech translation,”
Chengyi Wang et al., · 2020
Closest in time.
“Curriculum pre-training for end-to-end speech translation,”
Chengyi Wang et al., · 2020
Closest in time.
“Self-training for end-to-end speech translation,”
Juan Pino et al., · 2020
Closest in time.
“Phone features improve speech translation,”
Elizabeth Salesky et al., · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ye Jia et al., · 2019
Cited alongside, same era.
“Harnessing indirect training data for end-to-end automatic speech translation: Tricks of the trade,”
Juan Pino et al., · 2019
Cited alongside, same era.
“Multilingual end-to-end speech translation,”
Hirofumi Inaguma et al., · 2019
Cited alongside, same era.
“One-to-many multilingual end-to-end speech translation,”
Mattia Antonino Di Gangi et al., · 2019
Cited alongside, same era.
“Exploring phoneme-level speech representations for end-to-end speech translation,”
Elizabeth Salesky et al., · 2019
Cited alongside, same era.
“Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,”
Sameer Bansal et al., · 2019
Cited alongside, same era.
“End-to-end speech translation with knowledge distillation,”
Yuchen Liu et al., · 2019
Cited alongside, same era.
Closest in time.
“End-end speech-to-text translation with modality agnostic meta-learning,”
Sathish Indurthi et al., · 2020
Closest in time.
“Aligned cross entropy for non-autoregressive machine translation,”
Marjan Ghazvininejad et al., · 2020
Closest in time.
“Non-autoregressive machine translation with latent alignments,”
Chitwan Saharia et al., · 2020
Closest in time.
“Semi-autoregressive training improves mask-predict decoding,”
Marjan Ghazvininejad et al., · 2020
Closest in time.
“Imputer: Sequence modelling via imputation and dynamic programming,”
William Chan et al., · 2020
Closest in time.
“Mask CTC: Non-autoregressive end-to-end ASR with CTC and mask predict,”
Yosuke Higuchi et al., · 2020
Closest in time.
“Insertion-based modeling for end-to-end automatic speech recognition,”
Yuya Fujita et al., · 2020
Closest in time.
“Listen attentively, and spell once: Whole sentence generation via a non-autoregressive architecture for low-latency speech recognition,”
Ye Bai et al., · 2020
Closest in time.
“Spike-triggered non-autoregressive Transformer for end-to-end speech recognition,”
Zhengkun Tian et al., · 2020
Closest in time.
“Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation,”
Junliang Guo et al., · 2020
Closest in time.
“A study of non-autoregressive model for sequence generation,”
Yi Ren et al., · 2020
Closest in time.
“ESPnet-ST: All-in-one speech translation toolkit,”
Hirofumi Inaguma et al., · 2020
Closest in time.