2018

AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Du, Jiayu, Na, Xingyu, Liu, Xuechen et al.

Understand

AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research.

  • It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR.
  • In AISHELL-2, 1000 hours of clean read-speech data from iOS is published, which is free for academic usage.
  • On top of AISHELL-2 corpus, an improved recipe is developed and released, containing key components for industrial applications, such as Chinese word segmentation, flexible vocabulary expension and phone set transformation etc.

Built on

  • P. Fung, S. Huang, and D. Graff. (2005) Hkust mandarin telephone speech, part 1. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2005S15

    2005

    Earlier work this paper cites.

  • S. Duanmu,

    2007

    Earlier work this paper cites.

  • J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in

    2009

    Earlier work this paper cites.

  • D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi Speech Recognition Toolkit,” in

    2011

    Earlier work this paper cites.

  • N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”

    2011

    Earlier work this paper cites.

Similar

Then

  • H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AIShell-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline,” in

    2017

    Later among the works it cites.

  • V. H. Do, N. F. Chen, B. P. Lim, and M. A. Hasegawa-Johnson, “Multitask Learning for Phone Recognition of Underresourced Languages Using Mismatched Transcription,”

    2018

    Closest in time.

  • J. Li, X. Wang, Y. Zhao, and Y. Li, “Gated recurrent unit based acoustic modeling with future context,”

    Original

    2018

    Closest in time.

  • Y. Zhang, P. Zhang, and Y. Yan, “Data augmentation for language models via adversarial training,”

    2018

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…