Fetching the paper…
Reading the bibliography…
This paper proposes a novel online speaker diarization algorithm based on a fully supervised self-attention mechanism (SA-EEND).
“2000 NIST Speaker Recognition Evaluation,” https://catalog.ldc.upenn.edu/LDC2001S97
2000
Earlier work this paper cites.
“The NIST 1999 speaker recognition evaluation—an overview,”
Alvin Martin and Mark Przybocki, · 2000
Earlier work this paper cites.
“Corpus of Spontaneous Japanese: Its design and evaluation,”
Kikuo Maekawa, · 2003
Earlier work this paper cites.
“A spectral clustering approach to speaker diarization,”
Huazhong Ning, Ming Liu, Hao Tang, and Thomas S Huang, · 2006
Earlier work this paper cites.
“Probabilistic linear discriminant analysis,”
Sergey Ioffe, · 2006
Earlier work this paper cites.
“JHU ASPIRE system: Robust LVCSR with TDNNs, iVector adaptation and RNN-LMS,”
Vijayaditya Peddinti, Guoguo Chen, Vimal Manohar, Tom Ko, Daniel Povey, and Sanjeev Khudanpur, · 2006
Earlier work this paper cites.
“Acoustic beamforming for speaker diarization of meetings,”
Xavier Anguera, Chuck Wooters, and Javier Hernando, · 2007
Earlier work this paper cites.
“Improved novelty detection for online gmm based speaker diarization,”
Konstantin Markov and Satoshi Nakamura, · 2008
Earlier work this paper cites.
“GMM-UBM based open-set online speaker diarization,”
Jürgen Geiger, Frank Wallhoff, and Gerhard Rigoll, · 2010
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Unsupervised methods for speaker diarization: An integrated and iterative approach,”
Stephen H Shum, Najim Dehak, Réda Dehak, and James R Glass, · 2013
Earlier work this paper cites.
“Speaker diarization with plda i-vector scoring and unsupervised calibration,”
Gregory Sell and Daniel Garcia-Romero, · 2014
Earlier work this paper cites.
“Integrating online i-vector extractor with information bottleneck based speaker diarization system,”
Srikanth Madikeri, Ivan Himawan, Petr Motlicek, and Marc Ferras, · 2015
Earlier work this paper cites.
“MUSAN: A music, speech, and noise corpus,”
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Cited alongside, same era.
“Online speaker diarization using adapted i-vector transforms,”
Weizhong Zhu and Jason Pelecanos, · 2016
Cited alongside, same era.
“Speaker diarization using deep neural network embeddings,”
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree, · 2017
Cited alongside, same era.
“Developing on-line speaker diarization system.,”
Dimitrios Dimitriadis and Petr Fousek, · 2017
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“Deep neural network embeddings for text-independent speaker verification.,”
“All-neural online source separation, counting, and diarization for meeting analysis,”
Thilo von Neumann, Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani, and Reinhold Haeb-Umbach, · 2019
Later among the works it cites.
“Fully supervised speaker diarization,”
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang, · 2019
Later among the works it cites.
“End-to-end neural speaker diarization with permutation-free objectives,”
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Kenji Nagamatsu, and Shinji Watanabe, · 2019
Later among the works it cites.
“End-to-end neural speaker diarization with self-attention,”
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Yawen Xue, Kenji Nagamatsu, and Shinji Watanabe, · 2019
Later among the works it cites.
“Short utterance compensation in speaker verification via cosine-based teacher-student learning of speaker embeddings,”
Jee-weon Jung, Hee-Soo Heo, Hye-jin Shim, and Ha-Jin Yu, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Snyder, Daniel Garcia-Romero, Daniel Povey, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“The fifth “CHiME” speech separation and recognition challenge: Dataset, task and baselines,”
Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal, · 2018
Cited alongside, same era.
“The Hitachi/JHU CHiME-5 system: Advances in speech recognition for everyday home environments using multiple microphone arrays,”
Naoyuki Kanda, Rintaro Ikeshita, Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu, Xiaofei Wang, Vimal Manohar, Nelson Enrique Yalta Soplin, Matthew Maciejewski, Szu-Jui Chen, et al., · 2018
Cited alongside, same era.
“Characterizing performance of speaker diarization systems on far-field speech using standard methods,”
Matthew Maciejewski, David Snyder, Vimal Manohar, Najim Dehak, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Speaker diarization with LSTM,”
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno, · 2018
Cited alongside, same era.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
“X-vectors: Robust DNN embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“A study of x-vector based speaker recognition on short utterances,”
Ahilan Kanagasundaram, Sridha Sridharan, Sriram Ganapathy, Prachi Singh, and Clinton B Fookes, · 2019
Later among the works it cites.
“Multimodal speaker diarization of real-world meetings using d-vectors with spatial features,”
Wonjune Kang, Brandon C Roy, and Wesley Chow, · 2020
Closest in time.
“CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
Shinji Watanabe, Michael Mandel, Jon Barker, and Emmanuel Vincent, · 2020
Closest in time.
“Supervised online diarization with sample mean loss for multi-domain data,”
Enrico Fini and Alessio Brutti, · 2020
Closest in time.
Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Yawen Xue, and Kenji Nagamatsu, · 2020
Closest in time.
“Neural speaker diarization with speaker-wise chain rule,”
Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Yawen Xue, Jing Shi, and Kenji Nagamatsu, · 2020
Closest in time.
“End-to-end speaker diarization for an unknown number of speakers with encoder-decoder based attractors,”
Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue, and Kenji Nagamatsu, · 2020
Closest in time.
“Optimal mapping loss: A faster loss for end-to-end speaker diarization,”
Qingjian Lin, Tingle Li, Lin Yang, Junjie Wang, and Ming Li, · 2020
Closest in time.