Fetching the paper…
Reading the bibliography…
Multilingual speech recognition for both monolingual and code-switching speech is a challenging task.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. Dahl, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE international conference on acoustics, speech and signal processing
2013
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Neural Information Processing Systems , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP , 2016
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 4835–4839
2017
Earlier work this paper cites.
T. Hori, S. Watanabe, Y. Zhang, and W. Chan, “Advances in joint ctc-attention based end-to-end speech recognition with a deep cnn encoder and rnn-lm,” Proc. Interspeech 2017 , pp. 949–953, 2017
2017
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in Oriental COCOSDA 2017 . IEEE, pp. 1–5
2017
Earlier work this paper cites.
A. Pratapa, “Language modeling for code-mixing: The role of linguistic theory based synthetic data,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
G. Lee, X. Yue, and H. Li, “Linguistically motivated parallel data augmentation for code-switch language modeling,” in Interspeech 2019
2019
Cited alongside, same era.
W. Hou, Y. Wang, S. Gao, and T. Shinozaki, “Meta-adapter: Efficient cross-lingual adaptation with meta-learning,” in ICASSP , 2021
2021
Later among the works it cites.
S. Dalmia, Y. Liu, S. Ronanki, and K. Kirchhoff, “Transformer-transducers for code-switched speech recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5859–5863
2021
Later among the works it cites.
N. Gaur, B. Farris, P. Haghani, I. Leal, P. J. Moreno, M. Prasad, B. Ramabhadran, and Y. Zhu, “Mixture of informed experts for multilingual speech recognition,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6234–6238
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ma, B. Ramabhadran, J. Emond, A. Rosenberg, and F. Biadsy, “Comparison of data augmentation and adaptation strategies for code-switched automatic speech recognition,” in ICASSP , 2019
2019
Cited alongside, same era.
A. Kannan, A. Datta, T. N. Sainath, E. Weinstein, B. Ramabhadran, Y. Wu, A. Bapna, Z. Chen, and S. Lee, “Large-scale multilingual speech recognition with a streaming end-to-end model,” Proc. Interspeech 2019 , pp. 2130–2134
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” Proc. Interspeech 2019 , pp. 2613–2617
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Lu, M. Huang, H. Li, J. Guo, and Y. Qian, “Bi-encoder transformer network for mandarin-english code-switching speech recognition using mixture of experts,” in Interspeech , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
G. Ma, P. Hu, J. Kang, S. Huang, and H. Huang, “Leveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech Recognition,” in Proc. Interspeech 2021 , pp. 306–310
2021
Cited alongside, same era.
C. Jacobs, Y. Matusevych, and H. Kamper, “Acoustic word embeddings for zero-resource languages using self-supervised contrastive learning and multilingual adaptation,” in 2021 IEEE Spoken Language Technology Workshop (SLT)
2021
Cited alongside, same era.
2021
Later among the works it cites.
D. Wang, S. Ye, X. Hu, S. Li, and X. Xu, “An end-to-end dialect identification system with transfer learning from a multilingual automatic speech recognition model.” in Interspeech , 2021, pp. 3266–3270
2021
Later among the works it cites.
G. Ma, P. Hu, N. Yolwas, S. Huang, and H. Huang, “PM-MMUT: Boosted Phone-mask Data Augmentation using Multi-Modeling Unit Training for Phonetic-Reduction-Robust E2E Speech Recognition,” in Proc. Interspeech 2022 , pp. 1021–1025
2022
Later among the works it cites.
B. Yan, C. Zhang, M. Yu, S.-X. Zhang, S. Dalmia, D. Berrebbi, C. Weng, S. Watanabe, and D. Yu, “Joint modeling of code-switched and monolingual asr via conditional factorization,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6412–6416
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Tian, J. Yu, C. Zhang, Y. Zou, and D. Yu, “LAE: Language-Aware Encoder for Monolingual and Multilingual ASR,” in Proc. Interspeech 2022 , 2022, pp. 3178–3182
2022
Later among the works it cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022
2022
Later among the works it cites.
Y. Kwon and S.-W. Chung, “Mole : Mixture of language experts for multi-lingual automatic speech recognition,” in ICASSP 2023 , 2023, pp. 1–5
2023
Closest in time.