Fetching the paper…
Reading the bibliography…
How to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data? Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, but neglect the representation discrepancy across modalities.
End-to-end speech translation with knowledge distillation
Yuchen Liu, Hao Xiong, Zhongjun He, Jiajun Zhang, Hua Wu, Haifeng Wang, and Chengqing Zong. 2019 · 1904
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Bridging the modality gap for speech-to-text translation
Yuchen Liu, Junnan Zhu, Jiajun Zhang, and Chengqing Zong. 2020b · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016 · 2016
Earlier work this paper cites.
An attentional model for speech translation without transcription
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, and Trevor Cohn. 2016 · 2016
Earlier work this paper cites.
OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
MuST-C: a Multilingual Speech Translation Corpus
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019a · 2017
Earlier work this paper cites.
Structured-based curriculum learning for end-to-end english-japanese speech translation
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Earlier work this paper cites.
Neural lattice-to-sequence models for uncertain inputs
Matthias Sperber, Graham Neubig, Jan Niehues, and Alex Waibel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017 · 2017
Earlier work this paper cites.
End-to-end automatic speech translation of audiobooks
Alexandre Berard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin. 2018 · 2018
Earlier work this paper cites.
Towards robust neural machine translation
Yong Cheng, Zhaopeng Tu, Fandong Meng, Junjie Zhai, and Yang Liu. 2018 · 2018
Earlier work this paper cites.
An investigation of mixup training strategies for acoustic models in asr
Ivan Medennikov, Yuri Y Khokhlov, Aleksei Romanenko, Dmitry Popov, Natalia A Tomashenko, Ivan Sorokin, and Alexander Zatvornitskiy. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Towards fluent translations from disfluent speech
Elizabeth Salesky, Susanne Burger, Jan Niehues, and Alex Waibel. 2018 · 2018
Earlier work this paper cites.
End-to-end speech translation with the transformer
Laura Cross Vila, Carlos Escolano, José AR Fonollosa, and Marta R Costa-jussà. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2018 · 2018
Cited alongside, same era.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2019 · 2019
Cited alongside, same era.
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Analyzing asr pretraining for low-resource speech-to-text translation
Mihaela C Stoian, Sameer Bansal, and Sharon Goldwater. 2020 · 2020
Later among the works it cites.
Mixup-transformer: Dynamic data augmentation for nlp tasks
Lichao Sun, Congying Xia, Wenpeng Yin, Tingting Liang, S Yu Philip, and Lifang He. 2020 · 2020
Later among the works it cites.
fairseq s2t: Fast speech-to-text modeling with fairseq
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020a · 2020
Later among the works it cites.
Adaptive feature selection for end-to-end speech translation
Biao Zhang, Ivan Titov, Barry Haddow, and Rico Sennrich. 2020 · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Learning shared semantic space for speech-to-text translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Juan Pino, Liezl Puzon, Jiatao Gu, Xutai Ma, Arya D McCarthy, and Deepak Gopinath. 2019 · 2019
Cited alongside, same era.
Fluent translations from disfluent speech in end-to-end speech translation
Elizabeth Salesky, Matthias Sperber, and Alexander Waibel. 2019 · 2019
Cited alongside, same era.
Self-attentional models for lattice inputs
Matthias Sperber, Graham Neubig, Ngoc-Quan Pham, and Alex Waibel. 2019 · 2019
Cited alongside, same era.
Manifold mixup: Better representations by interpolating hidden states
Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio. 2019 · 2019
Cited alongside, same era.
Effectively pretraining a speech translation decoder with machine translation data
Ashkan Alinejad and Anoop Sarkar. 2020 · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Mixtext: Linguistically-informed interpolation of hidden space for semi-supervised text classification
Jiaao Chen, Zichao Yang, and Diyi Yang. 2020 · 2020
Cited alongside, same era.
Chi Han, Mingxuan Wang, Heng Ji, and Lei Li. 2021 · 2021
Later among the works it cites.
AdaST: Dynamically adapting encoder states in the decoder for end-to-end speech-to-text translation
Wuwei Huang, Dexin Wang, and Deyi Xiong. 2021 · 2021
Later among the works it cites.
Source and target bidirectional knowledge distillation for end-to-end speech translation
Hirofumi Inaguma, Tatsuya Kawahara, and Shinji Watanabe. 2021a · 2021
Later among the works it cites.
Cascaded models with cyclic feedback for direct speech translation
Tsz Kin Lam, Shigehiko Schamoni, and Stefan Riezler. 2021b · 2021
Later among the works it cites.
Mixup decoding for diverse machine translation
Jicheng Li, Pengzhi Gao, Xuanfu Wu, Yang Feng, Zhongjun He, Hua Wu, and Haifeng Wang. 2021a · 2021
Later among the works it cites.
Mixspeech: Data augmentation for low-resource automatic speech recognition
Linghui Meng, Jin Xu, Xu Tan, Jindong Wang, Tao Qin, and Bo Xu. 2021 · 2021
Later among the works it cites.
Semantic data augmentation for end-to-end mandarin speech recognition
Jianwei Sun, Zhiyuan Tang, Heng Yin, Wei Wang, Xi Zhao, Shuaijiang Zhao, Xiaoning Lei, Wei Zou, and Xiangang Li. 2021 · 2021
Later among the works it cites.
A general multi-task learning framework to leverage text data for speech to text tasks
Yun Tang, Juan Pino, Changhan Wang, Xutai Ma, and Dmitriy Genzel. 2021b · 2021
Later among the works it cites.
Jointly trained transformers models for spoken language translation
Hari Krishna Vydana, Martin Karafiát, Katerina Zmolikova, Lukáš Burget, and Honza Černockỳ. 2021 · 2021
Later among the works it cites.
Stacked acoustic-and-textual encoding: Integrating the pre-trained models into speech translation encoders
Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Shen Huang, Qi Ju, Tong Xiao, and Jingbo Zhu. 2021 · 2021
Later among the works it cites.
End-to-end speech translation via cross-modal progressive training
Rong Ye, Mingxuan Wang, and Lei Li. 2021 · 2021
Later among the works it cites.
Neural machine translation with phrase-level universal visual representations
Qingkai Fang and Yang Feng. 2022 · 2022
Closest in time.
Prediction difference regularization against perturbation for neural machine translation
Dengji Guo, Zhengrui Ma, Min Zhang, and Yang Feng. 2022 · 2022
Closest in time.
Enhancing cross-lingual transfer by manifold mixup
Huiyun Yang, Huadong Chen, Hao Zhou, and Lei Li. 2022 · 2022
Closest in time.