Fetching the paper…
Reading the bibliography…
End-to-end speaker diarization for an unknown number of speakers is addressed in this paper.
“2000 NIST Speaker Recognition Evaluation,”
2000
Earlier work this paper cites.
K. Maekawa, “Corpus of spontaneous Japanese: Its design and evaluation,” in
2003
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Acoustic beamforming for speaker diarization of meetings,”
2007
Earlier work this paper cites.
X. Zhu, C. Barras, L. Lamel, and J.-L. Gauvain, “Multi-stage speaker diarization for conference and lecture meetings,” in
2007
Earlier work this paper cites.
P. Kenny, D. Reynolds, and F. Castaldo, “Diarization of telephone conversations using factor analysis,”
2010
Earlier work this paper cites.
F. Vallet, S. Essid, and J. Carrive, “A multimodal approach to speaker diarization on TV talk-shows,”
2012
Earlier work this paper cites.
S. H. Shum, N. Dehak, R. Dehak, and J. R. Glass, “Unsupervised methods for speaker diarization: An integrated and iterative approach,”
2013
Earlier work this paper cites.
G. Sell and D. Garcia-Romero, “Speaker diarization with PLDA i-vector scoring and unsupervised calibration,” in
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Earlier work this paper cites.
D. Snyder, G. Chen, and D. Povey, “MUSAN: A music, speech, and noise corpus,” arXiv:1510.08484, 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
V. Peddinti, G. Chen, V. Manohar, T. Ko, D. Povey, and S. Khudanpur, “JHU ASpIRE system: Robust LVCSR with TDNNs, iVector adaptation and RNN-LMS,” in
2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” in
2016
Earlier work this paper cites.
I. Kapsouras, A. Tefas, N. Nikolaidis, G. Peeters, L. Benaroya, and I. Pitas, “Multimodal speaker clustering in full length movies,”
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
Cited alongside, same era.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in
2017
Cited alongside, same era.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
Q. Wang, C. Downey, L. Wan, P. Andrew Mansfield, and I. Lopez Moreno, “Speaker diarization with LSTM,” in
2018
Cited alongside, same era.
M. Diez, L. Burget, S. Wang, J. Rohdin, and J. Černockỳ, “Bayesian HMM based x-vector clustering for speaker diarization,” in
2019
Later among the works it cites.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in
2019
Later among the works it cites.
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with permutation-free objectives,” in
2019
Later among the works it cites.
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with self-attention,” in
2019
Later among the works it cites.
T. von Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, and R. Haeb-Umback, “All-neural online source separation, counting, and diarization for meeting analysis,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, “Listening to each speaker one by one with recurrent selective hearing networks,” in
2018
Cited alongside, same era.
J. Shi, J. Xu, G. Liu, and B. Xu, “Listen, think, and listen again: Captureing top-down auditory attention for speaker-independent speech separation,” in
2018
Cited alongside, same era.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,”
2018
Cited alongside, same era.
B. B. Meier, I. Elezi, M. Amirian, O. Dürr, and T. Stadelmann, “Learning neural models for end-to-end clustering,” in
2018
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in
2018
Cited alongside, same era.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in
2018
Cited alongside, same era.
N. Kanda, C. Boeddeker, J. Heitkaemper, Y. Fujita, S. Horiguchi, K. Nagamatsu, and R. Haeb-Umbach, “Guided source separation meets a strong ASR backend: Hitachi/Paderborn University joint investigation for dinner party scenario,” in
2019
Cited alongside, same era.
N. Takahashi, S. Parthasaarathy, N. Goswami, and Y. Mitsufuji, “Recursive speech separation for unknown number of speakers,” in
2019
Later among the works it cites.
J. Lee, Y. Lee, J. Kim, A. R. Kosiorek, S. Choi, and Y. W. Teh, “Set Transformer: A framework for attention-based permutation-invariant neural networks,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, “The Second DIHARD Diarization Challenge: Dataset, task, and baselines,” in
2019
Later among the works it cites.
S. Novoselov, A. Gusev, A. Ivanov, T. Pekhovsky, A. Shulipa, A. Avdeeva, A. Gorlanov, and A. Kozlov, “Speaker diarization with deep speaker embeddings for DIHARD Challenge II,” in
2019
Later among the works it cites.
Z. Huang, S. Watanabe, Y. Fujita, P. García, Y. Shao, D. Povey, and S. Khudanpur, “Speaker diarization with region proposal network,” in
2020
Closest in time.
2020
Closest in time.
F. Landini, S. Wang, M. Diez, L. Burget, P. Matějka, K. Žmolíková, L. Mošner, A. Silnova, O. Plchot, O. Novotnỳ, H. Zeinali, and S. Rohdin, “BUT system for the Second DIHARD Speech Diarization Challenge,” in
2020
Closest in time.