Fetching the paper…
Reading the bibliography…
End-to-end models with large capacity have significantly improved multilingual automatic speech recognition, but their computation cost poses challenges for on-device applications.
2012
Earlier work this paper cites.
S. Watanabe, T. Hori, and J. R. Hershey, “Language independent end-to-end architecture for joint language identification and speech recognition,” in 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2017, pp. 265–271
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Kim and M. L. Seltzer, “Towards language-universal end-to-end speech recognition,” in ICASSP . IEEE, 2018, pp. 4914–4918
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
W. Hou, Y. Dong, B. Zhuang, L. Yang, J. Shi, and T. Shinozaki, “Large-scale end-to-end multilingual speech recognition and language identification with multi-task learning,” in Interspeech , 2020, pp. 1037–1041
2020
Earlier work this paper cites.
Y. Zhu, P. Haghani, A. Tripathi, B. Ramabhadran, B. Farris, H. Xu, H. Lu, H. Sak, I. Leal, N. Gaur et al. , “Multilingual speech recognition with self-attention structured parameterization.” in INTERSPEECH , 2020, pp. 4741–4745
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Lu, M. Huang, H. Li, J. Guo, and Y. Qian, “Bi-encoder transformer network for mandarin-english code-switching speech recognition using mixture of experts.” in Interspeech , 2020, pp. 4766–4770
2020
Earlier work this paper cites.
X. Wang, F. Yu, L. Dunlap, Y.-A. Ma, R. Wang, A. Mirhoseini, T. Darrell, and J. E. Gonzalez, “Deep mixture of experts via shallow embedding,” in Uncertainty in artificial intelligence . PMLR, 2020, pp. 552–562
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in IEEE ICASSP , 2020, pp. 6139–6143
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. You, S. Feng, D. Su, and D. Yu, “SpeechMoE2: Mixture-of-experts model with improved routing,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7217–7221
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Gaur, B. Farris, P. Haghani, I. Leal, P. J. Moreno, M. Prasad, B. Ramabhadran, and Y. Zhu, “Mixture of informed experts for multilingual speech recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6234–6238
2021
Cited alongside, same era.
2021
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” J. Mach. Learn. Res , vol. 23, pp. 1–40, 2021
2021
Cited alongside, same era.
A. Narayanan, T. N. Sainath, R. Pang, J. Yu, C.-C. Chiu, R. Prabhavalkar, E. Variani, and T. Strohman, “Cascaded encoders for unifying streaming and non-streaming ASR,” in IEEE ICASSP , 2021, pp. 5629–5633
2021
Cited alongside, same era.
L. Zhou, J. Li, E. Sun, and S. Liu, “A configurable multilingual model is all you need to recognize all languages,” in ICASSP . IEEE, 2022, pp. 6422–6426
2022
Cited alongside, same era.
B. Li, T. N. Sainath, R. Pang, S.-y. Chang, Q. Xu, T. Strohman, V. Chen, Q. Liang, H. Liu, Y. He, P. Haghani, and S. Bidichandani, “A language agnostic multilingual streaming on-device ASR system,” in Interspeech , 2022, pp. 3188–3192
2022
Cited alongside, same era.
B. Li, R. Pang, Y. Zhang, T. N. Sainath, T. Strohman, P. Haghani, Y. Zhu, B. Farris, N. Gaur, and M. Prasad, “Massively multilingual ASR: A lifelong learning solution,” in ICASSP . IEEE, 2022, pp. 6397–6401
2022
Cited alongside, same era.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat et al. , “GLaM: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning . PMLR, 2022, pp. 5547–5569
2022
Later among the works it cites.
“Artificial intelligence at google: Our principles.” https://ai.google/principles/ , accessed: 2022-07-20
2022
Later among the works it cites.
K. Hu, B. Li, and T. N. Sainath, “Scaling up deliberation for multilingual ASR,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 771–776
2023
Closest in time.
2023
Closest in time.
S. Mavandadi, B. Li, C. Zhang, B. Farris, T. N. Sainath, and T. Strohman, “A truly multilingual first pass and monolingual second pass streaming on-device ASR system,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 838–845
2023
Closest in time.
2023
Closest in time.