Fetching the paper…
Reading the bibliography…
While recent research advances in speaker diarization mostly focus on improving the quality of diarization results, there is also an increasing interest in improving the efficiency of diarization systems.
“Multi-style training for robust isolated-word speech recognition,”
Richard Lippmann, Edward Martin, and D Paul, · 1987
Earlier work this paper cites.
“Matrix multiplication via arithmetic progressions,”
Don Coppersmith and Shmuel Winograd, · 1987
Earlier work this paper cites.
“TIMIT acoustic-phonetic continuous speech corpus LDC93S1,”
John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, David S. Pallett, Nancy L. Dahlgren, and Victor Zue, · 1993
Earlier work this paper cites.
“CALLHOME American English speech LDC97S42,” LDC Catalog. Philadelphia: Linguistic Data Consortium, 1997
A Canavan, D Graff, and G Zipperlen, · 1997
Earlier work this paper cites.
“The complexity of the matrix eigenproblem,”
Victor Y Pan and Zhao Q Chen, · 1999
Earlier work this paper cites.
“Efficient clustering of high-dimensional data sets with application to reference matching,”
Andrew McCallum, Kamal Nigam, and Lyle H Ungar, · 2000
Earlier work this paper cites.
“The ICSI meeting corpus,”
Adam Janin et al., · 2003
Earlier work this paper cites.
“The Fisher corpus: A resource for the next generations of speech-to-text,”
Christopher Cieri, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“Fisher English training speech parts 1 and 2, LDC2004S13, LDC2005S13,”
Christopher Cieri, David Graff, Owen Kimball, Dave Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
Jean Carletta et al., · 2005
Earlier work this paper cites.
“MapReduce: simplified data processing on large clusters,”
Jeffrey Dean and Sanjay Ghemawat, · 2008
Earlier work this paper cites.
“Multicondition training of Gaussian PLDA models in i-vector space for noise and reverberation robust speaker recognition,”
Daniel Garcia-Romero, Xinhui Zhou, and Carol Y. Espy-Wilson, · 2012
Earlier work this paper cites.
“Towards noise-robust speaker recognition using probabilistic linear discriminant analysis,”
Yun Lei, Lukas Burget, Luciana Ferrer, Martin Graciarena, and Nicolas Scheffer, · 2012
Earlier work this paper cites.
“Improving the performance of far-field speaker verification using multi-condition training: The case of GMM-UBM and i-vector systems,”
Anderson R. Avila, Milton Sarria-Paja, Francisco J. Fraga, Douglas O’Shaughnessy, and Tiago H. Falk, · 2014
Earlier work this paper cites.
“Automatic gain control and multi-style training for robust small-footprint keyword spotting with deep neural networks,”
Rohit Prabhavalkar, Raziel Alvarez, Carolina Parada, Preetum Nakkiran, and Tara N Sainath, · 2015
Earlier work this paper cites.
“Feature learning with raw-waveform CLDNNs for voice activity detection,”
Rubén Zazo Candil, Tara N Sainath, Gabor Simko, and Carolina Parada, · 2016
Earlier work this paper cites.
“Speaker diarization using deep neural network embeddings,”
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree, · 2017
Earlier work this paper cites.
“Developing on-line speaker diarization system,”
Dimitrios Dimitriadis and Petr Fousek, · 2017
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani, · 2017
Cited alongside, same era.
“ pyannote.metrics
Hervé Bredin, · 2017
Cited alongside, same era.
“Speaker diarization with LSTM,”
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
“Attention-based neural network for joint diarization and speaker extraction,”
“Transformer transducer: One model unifying streaming and non-streaming speech recognition,”
Anshuman Tripathi, Jaeyoung Kim, Qian Zhang, Han Lu, and Hasim Sak, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati et al., · 2020
Later among the works it cites.
“LibriVox: Free public domain audiobooks,” https://librivox.org/, 2020
LibriVox, · 2020
Later among the works it cites.
“CN-Celeb: A challenging Chinese speaker recognition dataset,”
Y. Fan, J.W. Kang, L.T. Li, K.C. Li, H.L. Chen, S.T. Cheng, P.Y. Zhang, Z.Y. Zhou, Y.Q. Cai, and D. Wang, · 2020
Later among the works it cites.
“Mixer 4 and 5 speech LDC2020S03,”
Linda Brandschain, Kevin Walker, David Graff, Christopher Cieri, Abby Neely, Nikki Mirghafori, Barbara Peskin, Jack Godfrey, Stephanie Strassel, Fred Goodman, George R. Doddington, and Mike King, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shlomo E Chazan, Sharon Gannot, and Jacob Goldberger, · 2018
Cited alongside, same era.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
“X-vectors: Robust DNN embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“First DIHARD challenge evaluation plan,”
Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristia, Jun Du, Sriram Ganapathy, and Mark Liberman, · 2018
Cited alongside, same era.
“Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,”
Tae Jin Park, Kyu J Han, Manoj Kumar, and Shrikanth Narayanan, · 2019
Cited alongside, same era.
“Fully supervised speaker diarization,”
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang, · 2019
Cited alongside, same era.
“End-to-end neural speaker diarization with permutation-free objectives,”
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Kenji Nagamatsu, and Shinji Watanabe, · 2019
Cited alongside, same era.
“Multi-scale speaker diarization with neural affinity score fusion,”
Tae Jin Park, Manoj Kumar, and Shrikanth Narayanan, · 2021
Later among the works it cites.
“Discriminative neural clustering for speaker diarisation,”
Qiujia Li, Florian L Kreyssig, Chao Zhang, and Philip C Woodland, · 2021
Later among the works it cites.
“Dr-Vectors: Decision residual networks and an improved loss for speaker recognition,”
Jason Pelecanos, Quan Wang, and Ignacio Lopez Moreno, · 2021
Later among the works it cites.
“SpeakerStew: Scaling to many languages with a triaged multilingual text-dependent and text-independent speaker verification system,”
Roza Chojnacka, Jason Pelecanos, Quan Wang, and Ignacio Lopez Moreno, · 2021
Later among the works it cites.
“Turn-to-Diarize: Online speaker diarization constrained by transformer transducer speaker turn detection,”
Wei Xia, Han Lu, Quan Wang, Anshuman Tripathi, Yiling Huang, Ignacio Lopez Moreno, and Hasim Sak, · 2022
Closest in time.
Yushi Ueda, Soumi Maiti, Shinji Watanabe, Chunlei Zhang, Meng Yu, Shi-Xiong Zhang, and Yong Xu, · 2022
Closest in time.
“A review of speaker diarization: Recent advances with deep learning,”
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu J Han, Shinji Watanabe, and Shrikanth Narayanan, · 2022
Closest in time.
“Speaker diarization: A journey from unsupervised to supervised approaches,” Odyssey: The Speaker and Language Recognition Workshop, 2022,
Chao Zhang and Quan Wang, · 2022
Closest in time.
“Attentive temporal pooling for conformer-based streaming language identification in long-form speech,”
Quan Wang, Yang Yu, Jason Pelecanos, Yiling Huang, and Ignacio Lopez Moreno, · 2022
Closest in time.
“Parameter-free attentive scoring for speaker verification,”
Jason Pelecanos, Quan Wang, Yiling Huang, and Ignacio Lopez Moreno, · 2022
Closest in time.
“Augmenting transformer-transducer based speaker change detection with token-level training loss,”
Guanlong Zhao, Quan Wang, Han Lu, Yiling Huang, and Ignacio Lopez Moreno, · 2023
Closest in time.
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding, Kohei Yatabe, Nobuyuki Morioka, Yu Zhang, Wei Han, Ankur Bapna, and Michiel Bacchiani, · 2023
Closest in time.