Fetching the paper…
Reading the bibliography…
Recent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , Pittsburgh, 2006
2006
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. E. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. ICASSP , Vancouver, 2013
2013
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. ICASSP , Shanghai, 2016
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP , New Orleans, 2017
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness et al. , “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences , vol. 114, no. 13, pp. 3521–3526, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
B. Li, T. N. Sainath, K. C. Sim, M. Bacchiani et al. , “Multi-dialect speech recognition with a single sequence-to-sequence model,” in Proc. ICASSP , Calgary, 2018
2018
Earlier work this paper cites.
A. Mallya, D. Davis, and S. Lazebnik, “Piggyback: Adapting a single network to multiple tasks by learning to mask weights,” in Proc. ECCV , Munich, 2018
2018
Earlier work this paper cites.
A. Waters, N. Gaur, P. Haghani, P. Moreno et al. , “Leveraging language ID in multilingual end-to-end speech recognition,” in Proc. ASRU , Sentosa, 2019
2019
Earlier work this paper cites.
A. Kannan, A. Datta, T. N. Sainath, E. Weinstein et al. , “Large-scale multilingual speech recognition with a streaming end-to-end model,” in Proc. Interspeech , Graz, 2019
2019
Earlier work this paper cites.
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan et al. , “Continual lifelong learning with neural networks: A review,” Neural Networks , vol. 113, pp. 54–71, 2019
2019
Earlier work this paper cites.
D. Rolnick, A. Ahuja, J. Schwarz, T. Lillicrap et al. , “Experience replay for continual learning,” in Proc. NeurIPS , Vancouver, 2019
2019
Earlier work this paper cites.
A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny, “Efficient lifelong learning with A-GEM,” in Proc. ICLR , New Orleans, 2019
2019
Earlier work this paper cites.
S. Hou, X. Pan, C. C. Loy, Z. Wang et al. , “Learning a unified classifier incrementally via rebalancing,” in Proc. CVPR , Long Beach, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve et al. , “MLS: A large-scale multilingual dataset for speech research,” in Proc. Interspeech , Shanghai, 2020
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler et al. , “Common Voice: A massively-multilingual speech corpus,” in Proc. ACL , Marseille, 2020
2020
Cited alongside, same era.
Y. Zhu, P. Haghani, A. Tripathi, B. Ramabhadran et al. , “Multilingual speech recognition with self-attention structured parameterization,” in Proc. Interspeech , Shanghai, 2020
L. Zhou, J. Li, E. Sun, and S. Liu, “A configurable multilingual model is all you need to recognize all languages,” in Proc. ICASSP , Singapore, 2022
2022
Later among the works it cites.
Y. Lu, M. Huang, X. Qu, P. Wei et al. , “Language adaptive cross-lingual speech representation learning with sparse sharing sub-networks,” in Proc. ICASSP , Singapore, 2022
2022
Later among the works it cites.
B. Li, R. Pang, Y. Zhang, T. N. Sainath et al. , “Massively multilingual ASR: A lifelong learning solution,” in Proc. ICASSP , Singapore, 2022
2022
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman et al. , “Robust speech recognition via large-scale weak supervision,” in Proc. ICML , Hawaii, 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
P. Buzzega, M. Boschini, A. Porrello, D. Abati et al. , “Dark experience for general continual learning: A strong, simple baseline,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
C. Wang, M. Riviere, A. Lee, A. Wu et al. , “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proc. ACL , Bangkok, 2021
2021
Cited alongside, same era.
N. Gaur, B. Farris, P. Haghani, I. Leal et al. , “Mixture of informed experts for multilingual speech recognition,” in Proc. ICASSP , Toronto, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Li et al. , “Recent advances in end-to-end automatic speech recognition,” APSIPA Transactions on Signal and Information Processing , vol. 11, no. 1, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Wang, G. Ma, Y. Li, and B. Du, “Language-routing mixture of experts for multilingual and code-switching speech recognition,” in Proc. Interspeech , Dublin, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Yang, A. Tjandra, C. Liu, D. Zhang et al. , “Learning ASR pathways: A sparse multilingual ASR model,” in Proc. ICASSP , Rhodes, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.