Fetching the paper…
Reading the bibliography…
With the emergence of large language models (LLMs), multimodal models based on LLMs have demonstrated significant potential.
Generating Synthetic Speech from SpokenVocab for Speech Translation
Zhao, J.; Haffari, G.; and Shareghi, E. 2023 · 1981
Earlier work this paper cites.
Bridging the Modality Gap for Speech-to-Text Translation
Liu, Y.; Zhu, J.; Zhang, J.; and Zong, C. 2020 · 2010
Earlier work this paper cites.
Fairseq S2T: Fast speech-to-text modeling with fairseq
Wang, C.; Tang, Y.; Ma, X.; Wu, A.; Popuri, S.; Okhonko, D.; and Pino, J. 2020a · 2010
Earlier work this paper cites.
On knowledge distillation for direct speech translation
Gaido, M.; Di Gangi, M. A.; Negri, M.; and Turchi, M. 2020 · 2012
Earlier work this paper cites.
Librivox: Free public domain audiobooks
Kearns, J. 2014 · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V.; Chen, G.; Povey, D.; and Khudanpur, S. 2015 · 2015
Earlier work this paper cites.
Must-c: a multilingual speech translation corpus
Di Gangi, M. A.; Cattoni, R.; Bentivogli, L.; Negri, M.; and Turchi, M. 2019 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Sequence-to-Sequence Models Can Directly Translate Foreign Speech
Weiss, R. J.; Chorowski, J.; Jaitly, N.; Wu, Y.; and Chen, Z. 2017 · 2017
Earlier work this paper cites.
End-to-End Automatic Speech Translation of Audiobooks
Berard, A.; Besacier, L.; Kocabiyikoglu, A. C.; and Pietquin, O. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Post, M. 2018 · 2018
Earlier work this paper cites.
On Using SpecAugment for End-to-End Speech Translation
Bahar, P.; Zeyer, A.; Schlüter, R.; and Ney, H. 2019 · 2019
Earlier work this paper cites.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Bansal, S.; Kamper, H.; Livescu, K.; Lopez, A.; and Goldwater, S. 2019 · 2019
Earlier work this paper cites.
Data Efficient Direct Speechto-Text Translation with Modality Agnostic Meta-Learning
Indurthi, S.; Han, H.; Lakumarapu, N.; Lee, B.; Chung, I.; Kim, S.; and Kim, C. 2019 · 2019
Earlier work this paper cites.
End-to-End Speech Translation with Knowledge Distillation
Liu, Y.; Xiong, H.; Zhang, J.; He, Z.; Wu, H.; Wang, H.; and Zong, C. 2019 · 2019
Earlier work this paper cites.
Attention-Passing Models for Robust and Data-Efficient End-to-End Speech Translation
Sperber, M.; Neubig, G.; Niehues, J.; and Waibel, A. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020 · 2020
Earlier work this paper cites.
End-to-End Speech-Translation with Knowledge Distillation: FBK@IWSLT2020
Gaido, M.; Gangi, M. A. D.; Negri, M.; and Turchi, M. 2020 · 2020
Earlier work this paper cites.
Synchronous speech recognition and speech-to-text translation with interactive decoding
Liu, Y.; Zhang, J.; Xiong, H.; Zhou, L.; He, Z.; Wu, H.; Wang, H.; and Zong, C. 2020 · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S.; Rasley, J.; Ruwase, O.; and He, Y. 2020 · 2020
Cited alongside, same era.
Analyzing ASR pretraining for low-resource speech-to-text translation
Stoian, M. C.; Bansal, S.; and Goldwater, S. 2020 · 2020
Cited alongside, same era.
Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text Translation
Dong, Q.; Ye, R.; Wang, M.; Zhou, H.; Xu, S.; Xu, B.; and Li, L. 2021 · 2021
Cited alongside, same era.
Upc’s speech translation system for iwslt 2021
Gállego, G. I.; Tsiamas, I.; Escolano, C.; Fonollosa, J. A.; and R Costa-jussa, M. 2021 · 2021
Cited alongside, same era.
Linearly mapping from image to text space
Merullo, J.; Castricato, L.; Eickhoff, C.; and Pavlick, E. 2022 · 2022
Later among the works it cites.
WACO: word-aligned contrastive learning for speech translation
Ouyang, S.; Ye, R.; and Li, L. 2022 · 2022
Later among the works it cites.
Unified speech-text pre-training for speech translation and recognition
Tang, Y.; Gong, H.; Dong, N.; Wang, C.; Hsu, W.-N.; Gu, J.; Baevski, A.; Li, X.; Mohamed, A.; Auli, M.; et al. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R.; De Freitas, D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.-T.; Jin, A.; Bos, T.; Baker, L.; Du, Y.; et al. 2022 · 2022
Later among the works it cites.
Cross-modal Contrastive Learning for Speech Translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning Shared Semantic Space for Speech-to-Text Translation
Han, C.; Wang, M.; Ji, H.; and Li, L. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021 · 2021
Cited alongside, same era.
Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation
Inaguma, H.; Kawahara, T.; and Watanabe, S. 2021 · 2021
Cited alongside, same era.
Multilingual Speech Translation from Efficient Finetuning of Pretrained Models
Li, X.; Wang, C.; Tang, Y.; Tran, C.; Tang, Y.; Pino, J.; Baevski, A.; Conneau, A.; and Auli, M. 2021 · 2021
Cited alongside, same era.
Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task
Tang, Y.; Pino, J.; Li, X.; Wang, C.; and Genzel, D. 2021 · 2021
Cited alongside, same era.
Self-training and pre-training are complementary for speech recognition
Xu, Q.; Baevski, A.; Likhomanenko, T.; Tomasello, P.; Conneau, A.; Collobert, R.; Synnaeve, G.; and Auli, M. 2021b · 2021
Cited alongside, same era.
Mutual-Learning Improves End-to-End Speech Translation
Zhao, J.; Luo, W.; Chen, B.; and Gilman, A. 2021 · 2021
Cited alongside, same era.
Ye, R.; Wang, M.; and Li, L. 2022 · 2022
Later among the works it cites.
Revisiting end-to-end speech-to-text translation from scratch
Zhang, B.; Haddow, B.; and Sennrich, R. 2022 · 2022
Later among the works it cites.
SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
Zhang, Z.; Zhou, L.; Ao, J.; Liu, S.; Dai, L.; Li, J.; and Wei, F. 2022 · 2022
Later among the works it cites.
Chen, F.; Han, M.; Zhao, H.; Zhang, Q.; Shi, J.; Xu, S.; and Xu, B. 2023 · 2023
Closest in time.
M 3 st: Mix at three levels for speech translation
Cheng, X.; Dong, Q.; Yue, F.; Ko, T.; Wang, M.; and Zou, Y. 2023 · 2023
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; Zhou, A.; Zhang, W.; Lu, P.; He, C.; Yue, X.; et al. 2023 · 2023
Closest in time.
Gong, Y.; Luo, H.; Liu, A. H.; Karlinsky, L.; and Glass, J. 2023 · 2023
Closest in time.
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Closest in time.
LLaSM: Large Language and Speech Model
Yu Shu, G. C. W. H. R. Z. D. S. . Y. S., Siwei Dong. 2023 · 2023
Closest in time.
Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech Translation
Zhang, H.; Si, N.; Chen, Y.; Zhang, W.; Yang, X.; Qu, D.; and Li, Z. 2023c · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.