Understand
AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research.
- It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR.
- In AISHELL-2, 1000 hours of clean read-speech data from iOS is published, which is free for academic usage.
- On top of AISHELL-2 corpus, an improved recipe is developed and released, containing key components for industrial applications, such as Chinese word segmentation, flexible vocabulary expension and phone set transformation etc.
Built on
P. Fung, S. Huang, and D. Graff. (2005) Hkust mandarin telephone speech, part 1. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2005S15
2005
Earlier work this paper cites.
S. Duanmu,
2007
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in
2009
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi Speech Recognition Toolkit,” in
2011
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
Similar
J. Sun, “Jieba Chinese word segmentation tool,” 2012
2012
Cited alongside, same era.
2014
Cited alongside, same era.
D. Wang and X. Zhang, “THCHS-30 : A free chinese speech corpus,”
2015
Cited alongside, same era.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for asr based on lattice-free mmi,” in
2016
Cited alongside, same era.
Then
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AIShell-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline,” in
2017
Later among the works it cites.
V. H. Do, N. F. Chen, B. P. Lim, and M. A. Hasegawa-Johnson, “Multitask Learning for Phone Recognition of Underresourced Languages Using Mismatched Transcription,”
2018
Closest in time.
2018
Closest in time.
Y. Zhang, P. Zhang, and Y. Yan, “Data augmentation for language models via adversarial training,”
2018
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…