Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks.
Breaking the data barrier: Towards robust speech translation via adversarial stability training
Qiao Cheng, Meiyuan Fang, Yaqian Han, Jin Huang, and Yitao Duan. 2019 · 1909
Earlier work this paper cites.
Espnet-st: All-in-one speech translation toolkit
Hirofumi Inaguma, Shun Kiyono, Kevin Duh, Shigeki Karita, Nelson Enrique Yalta Soplin, Tomoki Hayashi, and Shinji Watanabe. 2020 · 2004
Earlier work this paper cites.
From wer and ril to mer and wil: improved evaluation measures for connected speech recognition
Andrew Cameron Morris, Viktoria Maier, and Phil Green. 2004 · 2004
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al. 2020 · 2005
Earlier work this paper cites.
Improving cross-lingual transfer learning for end-to-end speech recognition with speech translation
Changhan Wang, Juan Pino, and Jiatao Gu. 2020a · 2006
Earlier work this paper cites.
Covost 2 and massively multilingual speech-to-text translation
Changhan Wang, Anne Wu, and Juan Pino. 2020c · 2007
Earlier work this paper cites.
Fairseq s2t: Fast speech-to-text modeling with fairseq
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Sravya Popuri, Dmytro Okhonko, and Juan Pino. 2020b · 2010
Earlier work this paper cites.
Must-c: a multilingual speech translation corpus
Mattia A Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019 · 2017
Earlier work this paper cites.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Neural speech translation using lattice transformations and graph networks
Daniel Beck, Trevor Cohn, and Gholamreza Haffari. 2019 · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber. 2020 · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Speech translation and the end-to-end promise: Taking stock of where we are
Matthias Sperber and Matthias Paulik. 2020 · 2020
Cited alongside, same era.
Consistent transcription and translation of speech
Matthias Sperber, Hendra Setiawan, Christian Gollan, Udhyakumar Nallasamy, and Matthias Paulik. 2020 · 2020
Cited alongside, same era.
Consecutive decoding for speech-to-text translation
Qianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, and Lei Li. 2021 · 2021
Cited alongside, same era.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Cited alongside, same era.
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023 · 2023
Later among the works it cites.
Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023 · 2023
Later among the works it cites.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. 2024 · 2024
Closest in time.
An embarrassingly simple approach for llm with strong asr capacity
Ziyang Ma, Guanrou Yang, Yifan Yang, Zhifu Gao, Jiaming Wang, Zhihao Du, Fan Yu, Qian Chen, Siqi Zheng, Shiliang Zhang, et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Fleurs: Few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. 2022 · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Cited alongside, same era.
Seamlessm4t-massively multilingual & multimodal machine translation
Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023 · 2023
Cited alongside, same era.
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
Pre-training for speech translation: Ctc meets optimal transport
Phuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino, Benjamin Lecouteux, and Didier Schwab. 2023 · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Cited alongside, same era.
Multi-task transfer matters during instruction-tuning
David Mueller, Mark Dredze, and Nicholas Andrews. 2024 · 2024
Closest in time.
Leveraging timestamp information for serialized joint streaming recognition and translation
Sara Papi, Peidong Wang, Junkun Chen, Jian Xue, Naoyuki Kanda, Jinyu Li, and Yashesh Gaur. 2024 · 2024
Closest in time.
Unibridge: A unified approach to cross-lingual transfer learning for low-resource languages
Trinh Pham, Khoi M Le, and Luu Anh Tuan. 2024 · 2024
Closest in time.
Blsp-kd: Bootstrapping language-speech pre-training via knowledge distillation
Chen Wang, Minpeng Liao, Zhongqiang Huang, and Jiajun Zhang. 2024 · 2024
Closest in time.
Connecting speech encoder and large language model for asr
Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2024 · 2024
Closest in time.
Perception, reason, think, and plan: A survey on large multimodal reasoning models
Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al. 2025 · 2025
Closest in time.
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025 · 2025
Closest in time.