Fetching the paper…
Reading the bibliography…
Most recent speech recognition models rely on large supervised datasets, which are unavailable for many low-resource languages.
CMU, “The CMU pronunciation dictionary,” 2000. [Online]. Available: http://www.speech.cs.cmu.edu
2000
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
K. P. Scannell, “The crubadan project: Corpus building for under-resourced languages,” Cahiers du Cental , vol. 5, p. 1, 2007
2007
Earlier work this paper cites.
S. Nordhoff and H. Hammarström, “Glottolog/langdoc: Defining dialects, languages, and language families as collections of resources,” in First International Workshop on Linked Science 2011-In conjunction with the International Semantic Web Conference (ISWC 2011) , 2011
2011
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE international conference on acoustics, speech and signal processing . Ieee, 2013, pp. 6645–6649
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep rnn models and wfst-based decoding,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . IEEE, 2015, pp. 167–174
2015
Earlier work this paper cites.
M. P. Lewis, Ed., Ethnologue: Languages of the World . Dallas, TX, USA: SIL International, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Veselý, L. Burget, and J. Černocký, “Semi-Supervised DNN Training with Word Selection for ASR,” in Proc. Interspeech 2017 , 2017, pp. 3687–3691
2017
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N.-E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen et al. , “ESPnet: End-to-end speech processing toolkit,” Proc. Interspeech 2018 , pp. 2207–2211, 2018
2018
Earlier work this paper cites.
M. Artetxe, G. Labaka, and E. Agirre, “A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 789–798
2018
Cited alongside, same era.
D. R. Mortensen, S. Dalmia, and P. Littell, “Epitran: Precision g2p for many languages,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , 2018
2018
Cited alongside, same era.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang et al. , “A comparative study on transformer vs rnn in speech applications,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 449–456
2019
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” 2020
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Later among the works it cites.
J. Xu, X. Tan, Y. Ren, T. Qin, J. Li, S. Zhao, and T.-Y. Liu, “Lrspeech: Extremely low-resource speech synthesis and recognition,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 2802–2812
2020
Later among the works it cites.
X. Li, S. Dalmia, J. Li, M. Lee, P. Littell, J. Yao, A. Anastasopoulos, D. R. Mortensen, G. Neubig, A. W. Black et al. , “Universal phone recognition with a multilingual allophone system,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 8249–8253
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. W. Black, “CMU wilderness multilingual speech dataset,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5971–5975
2019
Cited alongside, same era.
X. Li, S. Dalmia, A. W. Black, and F. Metze, “Multilingual speech recognition with corpus relatedness sampling,” Proc. Interspeech 2019 , pp. 2120–2124, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Rosenberg, Y. Zhang, B. Ramabhadran, Y. Jia, P. Moreno, Y. Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in 2019 IEEE automatic speech recognition and understanding workshop (ASRU) . IEEE, 2019, pp. 996–1002
2019
Cited alongside, same era.
J. Chorowski, R. J. Weiss, S. Bengio, and A. Van Den Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM transactions on audio, speech, and language processing , vol. 27, no. 12, pp. 2041–2053, 2019
2019
Cited alongside, same era.
A. Tjandra, B. Sisman, M. Zhang, S. Sakti, H. Li, and S. Nakamura, “ \text
2019
Cited alongside, same era.
S. Moran and D. McCloy, Eds., PHOIBLE 2.0 . Jena: Max Planck Institute for the Science of Human History, 2019. [Online]. Available: https://phoible.org/
2019
Cited alongside, same era.
2020
Later among the works it cites.
D. R. Mortensen, X. Li, P. Littell, A. Michaud, S. Rijhwani, A. Anastasopoulos, A. W. Black, F. Metze, and G. Neubig, “Allovera: A multilingual allophone database,” in Proceedings of the 12th Language Resources and Evaluation Conference , 2020, pp. 5329–5336
2020
Later among the works it cites.
2020
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, “Unsupervised speech recognition,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
X. Li, J. Li, F. Metze, and W. B. Black, Alan, “Hierarchical phone recognition with compositional phonetics,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Proc. Interspeech 2021 , 2021, pp. 1194–1198
2021
Later among the works it cites.
X. Li, F. Metze, D. R. Mortensen, S. Watanabe, and A. W. Black, “Zero-shot learning for grapheme to phoneme conversion with language ensemble,” To be appearing at Findings of ACL , 2022
2022
Closest in time.