Fetching the paper…
Reading the bibliography…
In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on , vol. 1. IEEE Computer Society, 1992, pp. 517–520
1992
Earlier work this paper cites.
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke et al. , “The ICSI meeting corpus,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). , vol. 1. IEEE, 2003, pp. 1–5
2003
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The Fisher corpus: a resource for the next generations of speech-to-text.” in LREC , vol. 4, 2004, pp. 69–71
2004
Earlier work this paper cites.
Ö. Çetin and E. Shriberg, “Analysis of overlaps in meetings by dialog factors, hot spots, speakers, and collection site: Insights for automatic speech recognition,” in Ninth international conference on spoken language processing , 2006
2006
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in European Conference on Computer Vision . Springer, 2006, pp. 531–542
2006
Earlier work this paper cites.
D. Mostefa, N. Moreau, K. Choukri, G. Potamianos, S. M. Chu, A. Tyagi, J. R. Casas, J. Turmo, L. Cristoforetti, F. Tobia et al. , “The CHIL audiovisual corpus for lecture and meeting analysis inside smart rooms,” Language resources and evaluation , vol. 41, no. 3, pp. 389–407, 2007
2007
Earlier work this paper cites.
S. Renals, T. Hain, and H. Bourlard, “Interpretation of multiparty meetings the AMI and AMIDA projects,” in Hands-Free Speech Communication and Microphone Arrays . IEEE, 2008, pp. 115–118
2008
Earlier work this paper cites.
K. J. Han, S. Kim, and S. S. Narayanan, “Strategies to improve the robustness of agglomerative hierarchical clustering under data source variation for speaker diarization,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 16, no. 8, pp. 1590–1601, 2008
2008
Earlier work this paper cites.
A. Stupakov, E. Hanusa, D. Vijaywargi, D. Fox, and J. Bilmes, “The design and collection of COSINE, a multi-microphone in situ speech corpus recorded in noisy environments,” Computer Speech & Language , vol. 26, no. 1, pp. 52–66, 2012
2012
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 31–35
2016
Earlier work this paper cites.
P. Ghahremani, V. Manohar, D. Povey, and S. Khudanpur, “Acoustic modelling from the signal domain using CNNs.” in Interspeech , 2016, pp. 3434–3438
2016
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) . IEEE, 2017, pp. 1–5
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Cited alongside, same era.
2018
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, and J. Li, “Continuous speech separation: dataset and analysis,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7284–7288
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux, “WHAMR!: Noisy and reverberant single-channel speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 696–700
2020
Later among the works it cites.
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Cited alongside, same era.
T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dimitriadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang et al. , “Advances in online audio-visual meeting transcription,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 276–283
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4690–4699
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Later among the works it cites.
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech & Language , vol. 60, p. 101027, 2020
2020
Later among the works it cites.
Y. Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y. Cai, and D. Wang, “CN-CELEB: a challenging chinese speaker recognition dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7604–7608
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
N. Kanda, X. Chang, Y. Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “Investigation of end-to-end speaker-attributed asr for continuous multi-talker recordings,” in IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 809–816
2021
Closest in time.
X. Wang, N. Kanda, Y. Gaur, Z. Chen, Z. Meng, and T. Yoshioka, “Exploring end-to-end multi-channel asr with bias information for meeting transcription,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 833–840
2021
Closest in time.
C. Li, Y. Luo, C. Han, J. Li, T. Yoshioka, T. Zhou, M. Delcroix, K. Kinoshita, C. Boeddeker, Y. Qian et al. , “Dual-Path RNN for long recording speech separation,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 865–872
2021
Closest in time.