Fetching the paper…
Reading the bibliography…
In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total.
“An overview of the SPHINX speech recognition system,”
K-F Lee, H-W Hon, Raj Reddy, · 1990
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukáš Burget, Ondřej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlíček, Yanmin Qian, Petr Schwarz, Jan Silovský, Georg Stemmer, Karel Veselý, · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition,”
Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, Brian Kingsbury, · 2012
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
George E Dahl, Dong Yu, Li Deng, Alex Acero, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Jieba chinese word segmentation tool,” 2012
J Sun, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, Florian Metze, · 2015
Earlier work this paper cites.
“Deep speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, Oriol Vinyals, · 2016
Earlier work this paper cites.
“Detecting text in natural image with connectionist text proposal network,”
Zhi Tian, Weilin Huang, Tong He, Pan He, Yu Qiao, · 2016
Earlier work this paper cites.
“An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,”
Baoguang Shi, Xiang Bai, Cong Yao, · 2016
Cited alongside, same era.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, Shinji Watanabe, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, Illia Polosukhin, · 2017
Cited alongside, same era.
“Joint CTC/attention decoding for end-to-end speech recognition,”
Takaaki Hori, Shinji Watanabe, John R Hershey, · 2017
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, Ruoming Pang, · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli, · 2020
Later among the works it cites.
“Mls: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, Ronan Collobert, · 2020
Later among the works it cites.
“The ASRU 2019 mandarin-english code-switching speech recognition challenge: open datasets, tracks, methods and results,”
Xian Shi, Qiangze Feng, Lei Xie, · 2020
Later among the works it cites.
“Cascade rnn-transducer: Syllable based streaming on-device mandarin speech recognition with a syllable-to-character converter,”
Xiong Wang, Zhuoyuan Yao, Xian Shi, Lei Xie, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, Hao Zheng, · 2017
Cited alongside, same era.
“Speech-Transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, Bo Xu, · 2018
Cited alongside, same era.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, Tsubasa Ochiai, · 2018
Cited alongside, same era.
“Aishell-2: Transforming mandarin ASR research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, Hui Bu, · 2018
Cited alongside, same era.
“Semi-orthogonal low-rank matrix factorization for deep neural networks,”
Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, Hainan Xu, Mahsa Yarmohammadi, Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“A time-restricted self-attention layer for ASR,”
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, Michael Auli, · 2019
Cited alongside, same era.
“Efficient conformer with prob-sparse attention mechanism for end-to-endspeech recognition,”
Xiong Wang, Sining Sun, Lei Xie, Long Ma, · 2021
Closest in time.
“WeNet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, Xin Lei, · 2021
Closest in time.
“HuBERT: How much can a bad teacher benefit ASR pre-training?,”
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, Abdelrahman Mohamed, · 2021
Closest in time.
“Unsupervised speech recognition,”
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael Auli, · 2021
Closest in time.
“The people’s speech: A large-scale diverse english speech recognition dataset for commercial usage,”
Daniel Galvez, Greg Diamos, Juan Manuel Ciro Torres, Juan Felipe Cerón, Keith Achorn, Anjali Gopi, David Kanter, Max Lam, Mark Mazumder, Vijay Janapa Reddi, · 2021
Closest in time.
“Gigaspeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe, Shuaijiang Zhao, Wei Zou, Xiangang Li, Xuchen Yao, Yongqing Wang, Zhao You, Zhiyong Yan, · 2021
Closest in time.
“Unsupervised cross-lingual representation learning for speech recognition,”
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, Michael Auli, · 2021
Closest in time.
“Recent developments on ESPnet toolkit boosted by conformer,”
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, Yuekai Zhang, · 2021
Closest in time.