Fetching the paper…
Reading the bibliography…
In this paper, we propose a fully supervised speaker diarization approach, named unbounded interleaved-state recurrent neural networks (UIS-RNN).
“Speech understanding systems: Report of a steering committee,”
Mark F. Medress, Franklin S Cooper, Jim W. Forgie, CC Green, Dennis H. Klatt, Michael H. O’Malley, Edward P Neuburg, Allen Newell, DR Reddy, B Ritea, et al., · 1977
Earlier work this paper cites.
“The icsi meeting corpus,”
Adam Janin, Don Baron, Jane Edwards, Dan Ellis, David Gelbart, Nelson Morgan, Barbara Peskin, Thilo Pfau, Elizabeth Shriberg, Andreas Stolcke, et al., · 2003
Earlier work this paper cites.
“A spectral clustering approach to speaker diarization.,”
Huazhong Ning, Ming Liu, Hao Tang, and Thomas S Huang, · 2006
Earlier work this paper cites.
“Stream-based speaker segmentation using speaker factors and eigenvoices,”
Fabio Castaldo, Daniele Colibro, Emanuele Dalmasso, Pietro Laface, and Claudio Vair, · 2008
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
Vinod Nair and Geoffrey E Hinton, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“Distance dependent chinese restaurant processes,”
David M Blei and Peter I Frazier, · 2011
Earlier work this paper cites.
“Unsupervised methods for speaker diarization: An integrated and iterative approach,”
Stephen H Shum, Najim Dehak, Réda Dehak, and James R Glass, · 2013
Earlier work this paper cites.
“A study of the cosine distance-based mean shift for telephone speech diarization,”
Mohammed Senoussaoui, Patrick Kenny, Themos Stafylakis, and Pierre Dumouchel, · 2014
Earlier work this paper cites.
“Speaker diarization with plda i-vector scoring and unsupervised calibration,”
Gregory Sell and Daniel Garcia-Romero, · 2014
Cited alongside, same era.
“Learning phrase representations using rnn encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Diarization resegmentation in the factor analysis subspace,”
Gregory Sell and Daniel Garcia-Romero, · 2015
Cited alongside, same era.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“End-to-end text-dependent speaker verification,”
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer, · 2016
Cited alongside, same era.
“Voxceleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Later among the works it cites.
“ pyannote.metrics
Hervé Bredin, · 2017
Later among the works it cites.
“Speaker diarization with lstm,”
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno, · 2018
Closest in time.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Closest in time.
“Links: A high-dimensional online clustering method,”
Philip Andrew Mansfield, Quan Wang, Carlton Downey, Li Wan, and Ignacio Lopez Moreno, · 2018
Closest in time.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree, · 2017
Cited alongside, same era.
“Speaker diarization using convolutional neural network for statistics accumulation refinement,”
Zbynĕk Zajíc, Marek Hrúz, and Ludĕk Müller, · 2017
Cited alongside, same era.
“Developing on-line speaker diarization system,”
Dimitrios Dimitriadis and Petr Fousek, · 2017
Cited alongside, same era.
Ye Jia, Yu Zhang, Ron J Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, et al., · 2018
Closest in time.
“Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2018
Closest in time.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Closest in time.