Fetching the paper…
Reading the bibliography…
This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech.
N. Minematsu, M. Sekiguchi, and K. Hirose, “Automatic estimation of one’s age with his/her speech based upon acoustic modeling techniques of speakers,” in
2002
Earlier work this paper cites.
Y. Liu, P. Fung, Y. Yang
2006
Earlier work this paper cites.
T. Polyakova and A. Bonafonte, “Learning from errors in grapheme-to-phoneme conversion,” in
2006
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu
2012
Earlier work this paper cites.
Y. Lei, N. Scheffer, L. Ferrer
2014
Earlier work this paper cites.
K. Han, D. Yu, and I. Tashev, “Speech emotion recognition using deep neural network and extreme learning machine,” in
2014
Earlier work this paper cites.
M. H. Bahari, M. McLaren, H. V. hamme
2014
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk
2015
Earlier work this paper cites.
D. Wang and X. Zhang, “THCHS-30 : A free Chinese speech corpus,”
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren
2016
Cited alongside, same era.
S. Mirsamadi, E. Barsoum, and C. Zhang, “Automatic speech emotion recognition using recurrent neural networks with local attention,” in
2017
Cited alongside, same era.
H. Bu, J. Du, X. Na
2017
Cited alongside, same era.
D. B. China, “Chinese standard Mandarin speech corpus,”
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in
2017
Cited alongside, same era.
2019
Later among the works it cites.
T. Kaneko, H. Kameoka, K. Tanaka
2019
Later among the works it cites.
J. Yamagishi, C. Veaux, K. MacDonald
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Shen, R. Pang, R. J. Weiss
2018
Cited alongside, same era.
W. Zou, D. Jiang, S. Zhao
2018
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. J. Weiss
2018
Cited alongside, same era.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in
2018
Cited alongside, same era.
Later among the works it cites.
2020
Closest in time.
W. Li, D. Jiang, W. Zou
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
A. Nagrani, J. S. Chung, W. Xie
2020
Closest in time.
A. T. Liu, S. Yang, P. Chi
2020
Closest in time.