Fetching the paper…
Reading the bibliography…
In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way.
“2000 NIST speaker recognition evaluation,” https://catalog.ldc.upenn.edu/LDC2001S97
2000
Earlier work this paper cites.
“The rich transcription 2006 spring meeting recognition evaluation,”
J. G. Fiscus, J. Ajot, M. Michel, and J. S. Garofolo, · 2006
Earlier work this paper cites.
“Speaker diarization with PLDA i-vector scoring and unsupervised calibration,”
G. Sell and D. Garcia-Romero, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“Speaker diarization using deep neural network embeddings,”
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Xception: Deep learning with depthwise separable convolutions,”
F. Chollet, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Automatic speech recognition errors detection and correction: A review,”
R. Errattahi, A. El Hannani, and H. Ouahmane, · 2018
Cited alongside, same era.
“End-to-end neural speaker diarization with permutation-free objectives,”
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, · 2019
Cited alongside, same era.
“End-to-end neural speaker diarization with self-attention,”
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, · 2019
Cited alongside, same era.
“A spelling correction model for end-to-end speech recognition,”
J. Guo, T. N. Sainath, and R. J. Weiss, · 2019
Cited alongside, same era.
“Investigation of transformer based spelling correction model for CTC-based end-to-end Mandarin speech recognition.,”
“End-to-end spelling correction conditioned on acoustic feature for code-switching speech recognition.,”
S. Zhang, J. Yi, Z. Tian, Y. Bai, J. Tao, X. Liu, and Z. Wen, · 2021
Later among the works it cites.
“End-to-end speaker diarization as post-processing,”
S. Horiguchi, P. Garcia, Y. Fujita, S. Watanabe, and K. Nagamatsu, · 2021
Later among the works it cites.
“The third DIHARD diarization challenge,”
R. Neville, S. Prachi, K. Venkat, V. Rajat, C. Kenneth, C. Christopher, D. Jun, G. Sriram, and L. Mark, · 2021
Later among the works it cites.
“A review of speaker diarization: Recent advances with deep learning,”
T. J. Park, N. Kanda, D. Dimitriadis, K. J. Han, S. Watanabe, and S. Narayanan, · 2022
Later among the works it cites.
“Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks,”
F. Landini, J. Profant, M. Diez, and L. Burget, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zhang, M. Lei, and Z. Yan, · 2019
Cited alongside, same era.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Y. Luo and N. Mesgarani, · 2019
Cited alongside, same era.
“End-to-end speaker diarization for an unknown number of speakers with encoder-decoder based attractors,”
S. Horiguchi, Y. Fujita, S. Watanabe, Y. Xue, and K. Nagamatsu, · 2020
Cited alongside, same era.
“Correction of automatic speech recognition with transformer sequence-to-sequence model,”
Oleksii Hrinchuk, Mariya Popova, and Boris Ginsburg, · 2020
Cited alongside, same era.
“Fastcorrect: Fast error correction with edit alignment for automatic speech recognition,”
Y. Leng, X. Tan, L. Zhu, J. Xu, R. Luo, L. Liu, T. Qin, X. Li, E. Lin, and T.-Y. Liu, · 2021
Cited alongside, same era.
“Improving the Naturalness of Simulated Conversations for End-to-End Neural Diarization,”
Natsuo Yamashita, Shota Horiguchi, and Takeshi Homma, · 2022
Later among the works it cites.
“From simulated mixtures to simulated conversations as training data for end-to-end neural diarization,”
F. Landini, A. Lozano-Diez, M. Diez, and L. Burget, · 2022
Later among the works it cites.
“Neural diarization with non-autoregressive intermediate attractors,”
Y. Fujita, T. Komatsu, R. Scheibler, Y. Kida, and T. Ogawa, · 2023
Closest in time.
“Multi-speaker and wide-band simulated conversations as training data for end-to-end neural diarization,”
F. Landini, M. Diez, A. Lozano-Diez, and L. Burget, · 2023
Closest in time.
“Lexical speaker error correction: Leveraging language models for speaker diarization error correction,”
R. Paturi, S. Srinivasan, and X. Li, · 2023
Closest in time.