Fetching the paper…
Reading the bibliography…
The sparsely-gated Mixture of Experts (MoE) can magnify a network capacity with a little computational complexity.
T. Schultz and A. Waibel, “Language-independent and language-adaptive acoustic modeling for speech recognition,” Speech Communication , vol. 35, no. 1, pp. 31–52, 2001
2001
Earlier work this paper cites.
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , 2006
2006
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in Proc. ICML , 2012
2012
Earlier work this paper cites.
A. Ghoshal, P. Swietojanski, and S. Renals, “Multilingual training of deep neural networks,” in Proc. ICASSP , 2013
2013
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in Proc. Interspeech 2014 , 2014, pp. 338–342
2014
Earlier work this paper cites.
J. Cui, B. Kingsbury, B. Ramabhadran, A. Sethy, K. Audhkhasi, X. Cui, E. Kislal, L. Mangu, M. Nussbaum-Thom, M. Picheny, , Z. Tüske, P. Golik, R. Schlüter, H. Ney, M. J. F. Gales, K. M. Knill, A. Ragni, H. Wang, and P. Woodland, “Multilingual representations for low resource speech recognition and keyword search,” in Proc. ASRU . IEEE, 2015, pp. 259–266
2015
Earlier work this paper cites.
D. Chen and B. K.-W. Mak, “Multitask learning of deep neural networks for low-resource speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 7, pp. 1172–1183, 2015
2015
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Berlin, Germany: Association for Computational Linguistics, Aug. 2016, pp. 1715–1725. [Online]. Available: https://aclanthology.org/P16-1162
2016
Earlier work this paper cites.
S. Watanabe, T. Hori, and J. R. Hershey, “Language independent end-to-end architecture for joint language identification and speech recognition,” in Proc. ASRU . IEEE, 2017, pp. 265–271
2017
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in ICLR , 2017. [Online]. Available: https://openreview.net/pdf?id=B1ckMDqlg
2017
Earlier work this paper cites.
S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and K. Rao, “Multilingual speech recognition with a single end-to-end model,” in Proc. ICASSP . IEEE, 2018, pp. 4904–4908
2018
Earlier work this paper cites.
S. Kim and M. L. Seltzer, “Towards language-universal end-to-end speech recognition,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
A. Waters, N. Gaur, P. Haghani, P. Moreno, and Z. Qu, “Leveraging language id in multilingual end-to-end speech recognition,” in Proc. ASRU . IEEE, 2019, pp. 928–935
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. H. M. Wong, M. J. F. Gales, and Y. Wang, “General sequence teacher–student learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 11, pp. 1725–1736, 2019
2019
Cited alongside, same era.
N. Gaur, B. Farris, P. Haghani, I. Leal, P. J. Moreno, M. Prasad, B. Ramabhadran, and Y. Zhu, “Mixture of informed experts for multilingual speech recognition,” in Proc. ICASSP , 2021
2021
Closest in time.
L. Zhou, J. Li, E. Sun, and S. Liu, “A configurable multilingual model is all you need to recognize all languages,” in Proc. ASRU , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
H. Inaguma, Y. Gaur, L. Lu, J. Li, and Y. Gong, “Minimum latency training strategies for streaming sequence-to-sequence ASR,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathia, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
K. Kumatani, D. Dimitriadis, Y. Gaur, R. Gmyr, E. S. Eskimez, J. Li, and M. Zeng, “Sequence-level self-learning with multi-task learning framework,” in Proc. Interspeech , October 2020
2020
Cited alongside, same era.
C. Wang, Y. Wu, Y. Qian, K. Kumatani, S. Liu, F. Wei, M. Zeng, and X. Huang, “Unispeech: Unified speech representation learning with labeled and unlabeled data,” 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Closest in time.
S. Raghavan and K. Shubham, “Hybrid unsupervised and supervised multitask learning for speech recognition in low resource languages,” in Proc. Workshop on Machine Learning in Speech and Language Processing , 2021
2021
Closest in time.
W. Hou, Y. Wang, S. Gao, and T. Shinozaki, “Meta-adapter: Efficient cross-lingual adaptation with meta-learning,” in Proc. ICASSP , 2021
2021
Closest in time.
2021
Closest in time.
Z. You, S. Feng, D. Su, and D. Yu, “SpeechMoE: Scaling to large acoustic models with dynamic routing mixture of experts,” in Proc. Interspeech , 2021
2021
Closest in time.
——, “SpeechMoE2: Mixture-of-experts model with improved routing,” CoRR , vol. arXiv:2111.11831, 2021
2021
Closest in time.
M. Karimi, C. Liu, K. Kumatani, Y. Qian, T. Wu, and J. Wu, “Deploying self-supervised learning in the wild for hybrid automaticspeech recognition,” in Submitted to ICASSP 2022 , 2022
2022
Closest in time.