Fetching the paper…
Reading the bibliography…
This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based CSR corpus,” in Proc. 2nd International Conference on Spoken Language Processing (ICSLP 1992) , 1992, pp. 899–902
1992
Earlier work this paper cites.
J. Godfrey, E. Holliman, and J. McDaniel, “Switchboard: telephone speech corpus for research and development,” in Proc. ICASSP , vol. 1, 1992, pp. 517–520 vol.1
1992
Earlier work this paper cites.
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke et al. , “The icsi meeting corpus,” in Proc. ICASSP , vol. 1. IEEE, 2003, pp. I–I
2003
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: a resource for the next generations of speech-to-text,” in Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04) . Lisbon, Portugal: European Language Resources Association (ELRA), May 2004
2004
Earlier work this paper cites.
Y. Liu, P. Fung, Y. Yang, C. Cieri, S. Huang, and D. Graff, “Hkust/mts: A very large scale mandarin telephone speech corpus,” in International Symposium on Chinese Spoken Language Processing . Springer, 2006, pp. 724–735
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd International Conference on Machine Learning , ser. ICML ’06. Association for Computing Machinery, 2006, p. 369–376
2006
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in European Conference on Computer Vision . Springer, 2006, pp. 531–542
2006
Earlier work this paper cites.
D. Mostefa, N. Moreau, K. Choukri, G. Potamianos, S. M. Chu, A. Tyagi, J. R. Casas, J. Turmo, L. Cristoforetti, F. Tobia et al. , “The chil audiovisual corpus for lecture and meeting analysis inside smart rooms,” Language resources and evaluation , vol. 41, no. 3, pp. 389–407, 2007
2007
Earlier work this paper cites.
S. Renals, T. Hain, and H. Bourlard, “Recognition and understanding of meetings the ami and amida projects,” in 2007 IEEE Workshop on Automatic Speech Recognition Understanding (ASRU) , 2007, pp. 238–247
2007
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. CONF. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in International conference on machine learning . PMLR, 2014, pp. 1764–1772
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Proc. ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
D. Wang and X. Zhang, “Thchs-30 : A free chinese speech corpus,” ArXiv , vol. abs/1512.01882, 2015
2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. ICASSP . IEEE, 2016, pp. 4960–4964
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid ctc/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in INTERSPEECH , 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
H. Miao, G. Cheng, P. Zhang, and Y. Yan, “Online hybrid ctc/attention end-to-end automatic speech recognition architecture,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1452–1465, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) . IEEE, 2017, pp. 1–5
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP , 2017, pp. 4835–4839
2017
Cited alongside, same era.
2017
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” Proc. ICASSP , pp. 5884–5888, 2018
2018
Cited alongside, same era.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation,” in International conference on speech and computer . Springer, 2018, pp. 198–208
2018
Cited alongside, same era.
S. Kim and F. Metze, “Dialog-context aware end-to-end speech recognition,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 434–440
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Proceedings of Interspeech , 2018, pp. 2207–2211
2018
Cited alongside, same era.
F. Wang, J. Cheng, W. Liu, and H. Liu, “Additive margin softmax for face verification,” IEEE Signal Processing Letters , vol. 25, no. 7, pp. 926–930, 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
Y. Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y. Cai, and D. Wang, “Cn-celeb: a challenging chinese speaker recognition dataset,” in Proc. ICASSP . IEEE, 2020, pp. 7604–7608
2020
Later among the works it cites.
2021
Later among the works it cites.
G. Chen, S. Chai, G.-B. Wang, J. Du, W. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Y. Wang, Z. You, and Z. Yan, “Gigaspeech: An evolving, multi-domain asr corpus with 10, 000 hours of transcribed audio,” in Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Deng, G. Cheng, H. Miao, P. Zhang, and Y. Yan, “History utterance embedding transformer lm for speech recognition,” in Proc. ICASSP , 2021, pp. 5914–5918
2021
Later among the works it cites.
R. Yang, G. Cheng, H. Miao, T. Li, P. Zhang, and Y. Yan, “Keyword search using attention-based end-to-end asr and frame-synchronous phoneme alignments,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3202–3215, 2021
2021
Later among the works it cites.
F. Landini, J. Profant, M. Diez, and L. Burget, “Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implementation and analysis on standard tasks,” Computer Speech & Language , vol. 71, p. 101254, 2022
2022
Closest in time.