Fetching the paper…
Reading the bibliography…
We introduce TitaNet-LID, a compact end-to-end neural network for Spoken Language Identification (LID) that is based on the ContextNet architecture.
K. Walker and S. M. Strassel, “The RATS radio traffic collection system,” in Odyssey 2012: The Speaker and Language Recognition Workshop, Singapore, June 25-28, 2012 , H. Li, B. Ma, and K. Lee, Eds. ISCA, 2012, pp. 291–297. [Online]. Available: http://www.isca-speech.org/archive/odyssey_2012/od12_291.html
2012
Earlier work this paper cites.
H. Li, B. Ma, and K. A. Lee, “Spoken language recognition: From fundamentals to practice,” Proceedings of the IEEE , vol. 101, pp. 1136–1159, 05 2013
2013
Earlier work this paper cites.
I. Lopez-Moreno, J. Gonzalez-Dominguez, O. Plchot, D. Martinez, J. Gonzalez-Rodriguez, and P. Moreno, “Automatic language identification using deep neural networks,” in ICASSP , 2014
2014
Earlier work this paper cites.
J. Gonzalez-Dominguez, I. Lopez-Moreno, H. Sak, J. González-Rodríguez, and P. J. Moreno, “Automatic language identification using long short-term memory recurrent neural networks,” in INTERSPEECH , 2014
2014
Earlier work this paper cites.
S. Ganapathy, K. Han, S. Thomas, M. Omar, M. Van Segbroeck, and S. Narayanan, “Robust language identification using convolutional neural network features,” in INTERSPEECH , 2014
2014
Earlier work this paper cites.
M. Van Segbroeck, R. Travadi, and S. S. Narayanan, “Rapid language identification,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 7, pp. 1118–1129, 2015
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in ICASSP , 2017
2017
Earlier work this paper cites.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,” in Speaker Odyssey , 2018
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in IEEE CVPR , 2018
2018
Earlier work this paper cites.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-Attentive Speaker Embeddings for Text-Independent Speaker Verification,” in Proc. Interspeech , 2018
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in ICASSP . IEEE, 2018
2018
Earlier work this paper cites.
H. Mazzawi, X. Gonzalvo, A. Kracun, P. Sridhar, N. Subrahmanya, I. Lopez-Moreno, H.-J. Park, and P. Violette, “Improving keyword spotting and language identification via neural architecture search at scale.” in Interspeech , 2019, pp. 1278–1282
2019
Cited alongside, same era.
P. Shen, X. Lu, S. Li, and H. Kawai, “Interactive learning of teacher-student model for short utterance spoken language identification,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5981–5985
2019
Cited alongside, same era.
W. Cai, D. Cai, S. Huang, and M. Li, “Utterance-level end-to-end language identification using attention-based cnn-blstm,” in ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2019, pp. 5991–5995
2019
Cited alongside, same era.
X. Miao, I. McLoughlin, and Y. Yan, “A New Time-Frequency Attention Mechanism for TDNN and CNN-LSTM-TDNN, with Application to Language Identification,” in Proc. Interspeech , 2019
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. M. Pino, A. Baevski, A. Conneau, and M. Auli, “Xls-r: Self-supervised cross-lingual speech representation learning at scale,” in Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Liu, L. P. G. Perera, A. W. H. Khong, S. J. Styles, and S. Khudanpur, “Pho-lid: A unified model incorporating acoustic-phonetic and phonotactic information for language identification,” in Interspeech , 2022
2022
Closest in time.
P. Shen, X. Lu, and H. Kawai, “Transducer-based language embedding for spoken language identification,” in Interspeech , 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” Proc. Interspeech 2019 , pp. 2613–2617, 2019
2019
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” in 21st Annual conference of the International Speech Communication Association (INTERSPEECH 2020) . International Speech Communication Association (ISCA), 2020, pp. 3830–3834
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” Proc. Interspeech 2020 , pp. 5036–5040, 2020
2020
Cited alongside, same era.
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang, “Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,” in ICASSP , 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
J. Valk and T. Alumäe, “Voxlingua107: a dataset for spoken language recognition,” in IEEE SLT , 2021
2021
Cited alongside, same era.
M. H. Radfar, R. Barnwal, R. Swaminathan, F.-J. Chang, G. P. Strimel, N. Susanj, and A. Mouchtaris, “Convrnn-t: Convolutional augmented recurrent neural network transducers for streaming speech recognition,” in Interspeech , 2022
2022
Closest in time.
N. R. Koluguri, T. Park, and B. Ginsburg, “Titanet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context,” in ICASSP , 2022
2022
Closest in time.
A. Conneau, A. Bapna, Y. Zhang, M. Ma, P. von Platen, A. Lozhkov, C. Cherry, Y. Jia, C. Rivera, M. Kale, D. van Esch, V. Axelrod, S. Khanuja, J. Clark, O. Firat, M. Auli, S. Ruder, J. Riesa, and M. Johnson, “Xtreme-s: Evaluating cross-lingual speech representations,” in Interspeech , 2022
2022
Closest in time.
K. Kukk and T. Alumäe, “Improving language identification of accented speech,” in Interspeech , 2022
2022
Closest in time.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” 2022 IEEE Spoken Language Technology Workshop (SLT) , pp. 798–805, 2022
2022
Closest in time.
T. M. Bartley, F. Jia, K. C. Puvvada, S. Kriman, and B. Ginsburg, “Accidental learners: Spoken language identification in multilingual self-supervised models,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
W. Han, Z. Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu, “ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context,” in Proc. Interspeech 2020 , 2020, pp. 3610–3614. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2059
2059
Closest in time.