Fetching the paper…
Reading the bibliography…
Almost none of the 2,000+ languages spoken in Africa have widely available automatic speech recognition systems, and the required data is also only available for a few languages.
R. K. Herbert, “The sociohistory of clicks in Southern Bantu,” Anthropological linguistics , pp. 295–315, 1990
1990
Earlier work this paper cites.
D. Odden, “Tone: African languages,” The Handbook of Phonological Theory , vol. 1, pp. 444–75, 1995
1995
Earlier work this paper cites.
S. Bird, “Strategies for representing tone in African writing systems,” Written Language & Literacy , vol. 2, no. 1, pp. 1–44, 1999
1999
Earlier work this paper cites.
S. T. Abate, W. Menzel, and B. Tafila, “An Amharic speech corpus for large vocabulary continuous speech recognition,” in Ninth European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
T. Niesler, “Language-dependent state clustering for multilingual speech recognition in Afrikaans, South African English, Xhosa and Zulu,” in Multilingual Speech and Language Processing , 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd International Conference on Machine Learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
S. Amuda, H. Bořil, A. Sangwan, and J. H. Hansen, “Limited resource speech recognition for Nigerian English,” in ICASSP 2010 . IEEE, 2010, pp. 5090–5093
2010
Earlier work this paper cites.
J. Badenhorst, C. Van Heerden, M. Davel, and E. Barnard, “Collecting and evaluating speech recognition corpora for 11 South African languages,” Language resources and evaluation , vol. 45, no. 3, pp. 289–309, 2011
2011
Earlier work this paper cites.
T. Schlippe, E. G. K. Djomgang, N. T. Vu, S. Ochs, and T. Schultz, “Hausa large vocabulary continuous speech recognition,” in Spoken Language Technologies for Under-Resourced Languages , 2012
2012
Earlier work this paper cites.
D. Van Niekerk and E. Barnard, “Tone realisation in a Yorùbá speech recognition corpus,” 2012
2012
Earlier work this paper cites.
H. Gelas, L. Besacier, and F. Pellegrino, “Developments of Swahili resources for an automatic speech recognition system,” in Spoken Language Technologies for Under-Resourced Languages , 2012
2012
Earlier work this paper cites.
D. Henselmans, T. Niesler, and D. Van Leeuwen, “Baseline speech recognition of South African languages using Lwazi and AST,” Proc. PRASA, Johannesburg, South Africa , pp. 30–35, 2013
2013
Earlier work this paper cites.
A. Atanda, S. Yusof, and M. Hariharan, “Yorùbá automatic speech recognition: A review,” in Rural ICT Development (RICTD) International Conference , vol. 1, no. 1, 2013, pp. 116–121
2013
Earlier work this paper cites.
S. K. Kimutai, E. Milgo, and D. Gichoya, “Isolated Swahili words recognition using Sphinx4,” International Journal of Emerging Science and Engineering (IJESE) , vol. 2, no. 2, pp. 2319–6378, 2013
2013
Earlier work this paper cites.
E. Barnard, M. H. Davel, C. van Heerden, F. De Wet, and J. Badenhorst, “The NCHLT speech corpus of the South African languages,” in Workshop Spoken Language Technologies for Under-resourced Languages (SLTU) , 2014
2014
Earlier work this paper cites.
M. J. Harvilla and R. M. Stern, “Least squares signal declipping for robust speech recognition,” in INTERSPEECH , 2014
2014
Cited alongside, same era.
J. Cui, B. Kingsbury, B. Ramabhadran, A. Sethy, K. Audhkhasi, X. Cui, E. Kislal, L. Mangu, M. Nussbaum-Thom, M. Picheny et al. , “Multilingual representations for low resource speech recognition and keyword search,” in 2015 IEEE workshop on automatic speech recognition and understanding (ASRU) . IEEE, 2015, pp. 259–266
2015
Cited alongside, same era.
A. Das, P. Jyothi, and M. Hasegawa-Johnson, “Automatic Speech Recognition Using Probabilistic Transcriptions in Swahili, Amharic, and Dinka,” in INTERSPEECH , 2016, pp. 3524–3528
2016
Cited alongside, same era.
O. Adetunmbi, O. Obe, and J. Iyanda, “Development of Standard Yorùbá speech-to-text system using HTK,” International Journal of Speech Technology , vol. 19, no. 4, pp. 929–944, 2016
2016
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Later among the works it cites.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in ICASSP 2020 . IEEE, 2020, pp. 6139–6143
2020
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Gauthier, L. Besacier, S. Voisin, M. Melese, and U. P. Elingui, “Collecting resources in Sub-Saharan African languages for automatic speech recognition: a case study of Wolof,” in 10th Language Resources and Evaluation Conference (LREC 2016) , 2016
2016
Cited alongside, same era.
2017
Cited alongside, same era.
S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and K. Rao, “Multilingual speech recognition with a single end-to-end model,” in ICASSP 2018 . IEEE, 2018, pp. 4904–4908
2018
Cited alongside, same era.
S. Dalmia, R. Sanabria, F. Metze, and A. W. Black, “Sequence-based multi-lingual low resource speech recognition,” in ICASSP 2018 . IEEE, 2018, pp. 4909–4913
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang et al. , “Streaming end-to-end speech recognition for mobile devices,” in ICASSP 2019 . IEEE, 2019, pp. 6381–6385
2019
Cited alongside, same era.
2020
Cited alongside, same era.
A. Misra, D. Hwang, Z. Huo, S. Garg, N. Siddhartha, A. Narayanan, and K. C. Sim, “A comparison of supervised and unsupervised pre-training of end-to-end models,” in INTERSPEECH , vol. 2021, 2021, pp. 731–735
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
A. A. Boakye-Yiadom, M. Qin, and R. Jing, “Research of Automatic Speech Recognition of Asante-Twi Dialect For Translation,” in Proceedings of the 2021 5th International Conference on Electronic Information Technology and Computer Engineering , 2021, pp. 1086–1094
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Yu, C.-C. Chiu, B. Li, S.-y. Chang, T. N. Sainath, Y. He, A. Narayanan, W. Han, A. Gulati, Y. Wu et al. , “Fastemit: Low-latency streaming ASR with sequence-level emission regularization,” in ICASSP 2021 . IEEE, 2021, pp. 6004–6008
2021
Later among the works it cites.
T. N. Sainath, Y. R. He, A. Narayanan, R. Botros, R. Pang, D. J. Rybach, C. Allauzen, E. Variani, J. Qin, A. Gruenstein et al. , “An efficient streaming non-recurrent on-device end-to-end model with improvements to rare-word modeling,” in INTERSPEECH , 2021, pp. 1777–1781
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Aksënova, Z. Chen, C.-C. Chiu, D. van Esch, P. Golik, W. Han, L. King, B. Ramabhadran, A. Rosenberg, S. Schwartz, and G. Wang, “Accented speech recognition: benchmarking, pre-training, and diverse data,” 2022
2022
Closest in time.
U. A. Ibrahim, M. M. Boukar, and M. A. Suleiman, “Development of Hausa dataset: a baseline for speech recognition,” Data in Brief , vol. 40, 2022
2022
Closest in time.
D. van Esch, T. Lucassen, S. Ruder, I. Caswell, and C. Rivera, “Writing System and Speaker Metadata for 2,800+ Language Varieties,” in Proceedings of LREC , 2022
2022
Closest in time.