Fetching the paper…
Reading the bibliography…
Recent advancements in deep learning have significantly enhanced multilingual automatic speech recognition (ASR) due to the development of advanced model architectures and available large-scale multilingual datasets.
S. Dou, E. Zhou, Y. Liu, S. Gao, W. Shen, L. Xiong, Y. Zhou, X. Wang, Z. Xi, X. Fan et al. , “LoRAMoE: Alleviating world knowledge forgetting in large language models via MoE-style plugin,” in Proc. ACL , 2024, pp. 1932–1945
1945
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE.” Journal of Machine Learning Research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. IEEE ICASSP , 2013, pp. 6645–6649
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. IEEE ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. IEEE ICASSP , 2016, pp. 4960–4964
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. IEEE ICASSP . IEEE, 2017, pp. 4835–4839
2017
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in Proc. ICLR , 2017
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in Proc. O-COCOSDA , 2017, pp. 1–5
2017
Earlier work this paper cites.
A. Kannan, A. Datta, T. N. Sainath, E. Weinstein, B. Ramabhadran, Y. Wu, A. Bapna, Z. Chen, and S. Lee, “Large-scale multilingual speech recognition with a streaming end-to-end model,” in Proc. ISCA Interspeech , 2019, pp. 2130–2134
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Proc. ISCA Interspeech , 2020, pp. 2757–2761
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” in Proc. ACL , 2020, pp. 4218–4222
2020
Cited alongside, same era.
N. Gaur, B. Farris, P. Haghani, I. Leal, P. J. Moreno, M. Prasad, B. Ramabhadran, and Y. Zhu, “Mixture of informed experts for multilingual speech recognition,” in Proc. IEEE ICASSP , 2021, pp. 6234–6238
2021
Cited alongside, same era.
J. Li et al. , “Recent advances in end-to-end automatic speech recognition,” APSIPA Trans. on Signal and Information Processing , vol. 11, no. 1, 2022
2022
Cited alongside, same era.
V. Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi et al. , “Scaling speech technology to 1,000+ languages,” Journal of Machine Learning Research , vol. 25, no. 97, pp. 1–52, 2024
2024
Later among the works it cites.
S. Li, Y. You, X. Wang, K. Ding, and G. Wan, “Enhancing multilingual speech recognition through language prompt tuning and frame-level language adapter,” in Proc. IEEE ICASSP , 2024, pp. 10 941–10 945
2024
Later among the works it cites.
A. Piñeiro-Martín, C. García-Mateo, L. Docio-Fernandez, M. del Carmen López-Pérez, and G. Rehm, “Weighted cross-entropy for low-resource languages in multilingual speech recognition,” in Proc. ISCA Interspeech , 2024, pp. 1235–1239
2024
Later among the works it cites.
T. Xu, K. Huang, P. Guo, Y. Zhou, L. Huang, H. Xue, and L. Xie, “Towards rehearsal-free multilingual ASR: A LoRA-based case study on Whisper,” in Proc. ISCA Interspeech , 2024, pp. 2534–2538
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Zhou, J. Li, E. Sun, and S. Liu, “A configurable multilingual model is all you need to recognize all languages,” in Proc. IEEE ICASSP , 2022, pp. 6422–6426
2022
Cited alongside, same era.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
F. C. Salinas, K. Kumatani, R. Gmyr, L. Liu, and Y. Shi, “Knowledge distillation for mixture of experts models in speech recognition,” Microsoft Tech. Report, MSR-TR-2022-6, May 2022, Tech. Rep., 2022
2022
Cited alongside, same era.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in Proc. IEEE SLT , 2023, pp. 798–805
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML , 2023, pp. 28 492–28 518
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Wang, G. Ma, Y. Li, and B. Du, “Language-routing mixture of experts for multilingual and code-switching speech recognition,” in Proc. ISCA Interspeech , 2023, pp. 1389–1393
2023
Cited alongside, same era.
C. Y. Kwok, J. Q. Yip, and E. S. Chng, “Continual learning optimizations for auto-regressive decoder of multilingual ASR systems,” in Proc. ISCA Interspeech , 2024, pp. 1225–1229
2024
Later among the works it cites.
Y. Khassanov, Z. Chen, T. Chen, T. Y. Chong, W. Li, L. Lu, and Z. Ma, “Extending multilingual ASR to new languages using supplementary encoder and decoder components,” in Proc. IEEE ICASSP , 2024, pp. 10 586–10 590
2024
Later among the works it cites.
Z. Song, J. Zhuo, Y. Yang, Z. Ma, S. Zhang, and X. Chen, “LoRA-Whisper: Parameter-efficient and extensible multilingual ASR,” in Proc. ISCA Interspeech , 2024, pp. 3934–3938
2024
Later among the works it cites.
W. Liu, J. Hou, D. Yang, M. Cao, and T. Lee, “A parameter-efficient language extension framework for multilingual ASR,” in Proc. ISCA Interspeech , 2024, pp. 3929–3933
2024
Later among the works it cites.
2024
Later among the works it cites.
T. P. Ferraz, M. Z. Boito, C. Brun, and V. Nikoulina, “Multilingual Distilwhisper: Efficient distillation of multi-task speech models via language-specific experts,” in Proc. IEEE ICASSP , 2024, pp. 10 716–10 720
2024
Later among the works it cites.