Fetching the paper…
Reading the bibliography…
Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost, 2016
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen · 2017
Earlier work this paper cites.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Matt Post · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Earlier work this paper cites.
A comparative study on end-to-end speech to text translation
Parnia Bahar, Tobias Bieschke, and Hermann Ney · 2019
Earlier work this paper cites.
On using specaugment for end-to-end speech translation
Parnia Bahar, Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2019
Earlier work this paper cites.
Adapting transformer to end-to-end spoken language translation
Mattia A Di Gangi, Matteo Negri, and Marco Turchi · 2019
Earlier work this paper cites.
Enhancing transformer for end-to-end speech-to-text translation
Mattia Antonino Di Gangi, Matteo Negri, Roldano Cattoni, Roberto Dessi, and Marco Turchi · 2019
Earlier work this paper cites.
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu · 2019
Earlier work this paper cites.
End-to-end speech translation with knowledge distillation
Yuchen Liu, Hao Xiong, Zhongjun He, Jiajun Zhang, Hua Wu, Haifeng Wang, and Chengqing Zong · 2019
Earlier work this paper cites.
Fluent translations from disfluent speech in end-to-end speech translation
Elizabeth Salesky, Matthias Sperber, and Alex Waibel · 2019
Earlier work this paper cites.
Effectively pretraining a speech translation decoder with machine translation data
Ashkan Alinejad and Anoop Sarkar · 2020
Earlier work this paper cites.
Multilingual speech translation with efficient finetuning of pretrained models
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models, 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Speech translation and the end-to-end promise: Taking stock of where we are
Matthias Sperber and Matthias Paulik · 2020
Cited alongside, same era.
Analyzing asr pretraining for low-resource speech-to-text translation
Mihaela C Stoian, Sameer Bansal, and Sharon Goldwater · 2020
Cited alongside, same era.
Multilingual translation with extensible multilingual pretraining and finetuning
Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan · 2020
Cited alongside, same era.
Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Qi Ju, Tong Xiao, Jingbo Zhu, et al · 2021
Later among the works it cites.
End-to-end speech translation via cross-modal progressive training
Rong Ye, Mingxuan Wang, and Lei Li · 2021
Later among the works it cites.
Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation
Renjie Zheng, Junkun Chen, Mingbo Ma, and Liang Huang · 2021
Later among the works it cites.
mslam: Massively multilingual joint pre-training for speech and text
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau · 2022
Later among the works it cites.
Maestro: Matched speech text representations through modality matching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Changhan Wang, Anne Wu, and Juan Pino · 2020
Cited alongside, same era.
Bridging the gap between pre-training and fine-tuning for end-to-end speech translation
Chengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang, and Ming Zhou · 2020
Cited alongside, same era.
Curriculum pre-training for end-to-end speech translation
Chengyi Wang, Yu Wu, Shujie Liu, Ming Zhou, and Zhenglu Yang · 2020
Cited alongside, same era.
Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, et al · 2021
Cited alongside, same era.
Slam: A unified encoder for speech and language modeling via speech-text joint pre-training
Ankur Bapna, Yu-an Chung, Nan Wu, Anmol Gulati, Ye Jia, Jonathan H Clark, Melvin Johnson, Jason Riesa, Alexis Conneau, and Yu Zhang · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Source and target bidirectional knowledge distillation for end-to-end speech translation
Hirofumi Inaguma, Tatsuya Kawahara, and Shinji Watanabe · 2021
Cited alongside, same era.
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro Moreno, Ankur Bapna, and Heiga Zen · 2022
Later among the works it cites.
M3st: Mix at three levels for speech translation
Xuxin Cheng, Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, and Yuexian Zou · 2022
Later among the works it cites.
Stemm: Self-learning with speech-text manifold mixup for speech translation
Qingkai Fang, Rong Ye, Lei Li, Yang Feng, and Mingxuan Wang · 2022
Later among the works it cites.
Sample, translate, recombine: Leveraging audio alignments for data augmentation in end-to-end speech translation
Tsz Kin Lam, Shigehiko Schamoni, and Stefan Riezler · 2022
Later among the works it cites.
Waco: Word-aligned contrastive learning for speech translation
Siqi Ouyang, Rong Ye, and Lei Li · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Later among the works it cites.
Unified speech-text pre-training for speech translation and recognition
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, et al · 2022
Later among the works it cites.
A weakly-supervised streaming multilingual speech model with truly zero-shot capability
Jian Xue, Peidong Wang, Jinyu Li, and Eric Sun · 2022
Later among the works it cites.
Cross-modal contrastive learning for speech translation
Rong Ye, Mingxuan Wang, and Lei Li · 2022
Later among the works it cites.
Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Lirong Dai, Jinyu Li, and Furu Wei · 2022
Later among the works it cites.
Google usm: Scaling automatic speech recognition beyond 100 languages
Yu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, et al · 2023
Closest in time.