Fetching the paper…
Reading the bibliography…
Speaker diarization is a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify "who spoke when".
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,
M. Kolbæk, D. Yu, Z.-H. Tan, J. Jensen, · 1913
Earlier work this paper cites.
The hungarian method for the assignment problem,
H. W. Kuhn, · 1955
Earlier work this paper cites.
Remeeting – Deep insights to conversations,
A. Guo, A. Faria, J. Riedhammer, · 1965
Earlier work this paper cites.
Some methods for classification and analysis of multivariate observations,
J. MacQueen, et al., · 1967
Earlier work this paper cites.
Developing a speech activity detection system for the darpa rats program,
T. Ng, B. Zhang, L. Nguyen, S. Matsoukas, X. Zhou, N. Mesgarani, K. Veselỳ, P. Matějka, · 1972
Earlier work this paper cites.
Improved MVDR beamforming using single-channel mask prediction networks,
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, J. Le Roux, · 1985
Earlier work this paper cites.
Segregation of speakers for speech recognition and speaker identification,
H. Gish, M. . Siu, R. Rohlicek, · 1991
Earlier work this paper cites.
Segregation of speakers for speech recognition and speaker identification.,
H. Gish, M.-H. Siu, J. R. Rohlicek, · 1991
Earlier work this paper cites.
An unsupervised, sequential learning algorithm for segmentation for speech waveforms with multiple speakers,
M.-H. Siu, Y. George, H. Gish, · 1992
Earlier work this paper cites.
Gisting conversational speech,
J. R. Rohlicek, D. Ayuso, M. Bates, R. Bobrow, A. Boulanger, H. Gish, P. Jeanrenaud, M. Meteer, M. Siu, · 1992
Earlier work this paper cites.
Recognition of continuous broadcast news with multiple unknown speakers and environments,
U. Jain, M. A. Siegler, S.-J. Doh, E. Gouvea, J. Huerta, P. J. Moreno, B. Raj, R. M. Stern, · 1996
Earlier work this paper cites.
Speaker clustering and transformation for speaker adaptation in large-vocabulary speech recognition systems,
M. Padmanabhan, L. R. Bahl, D. Nahamoo, M. A. Picheny, · 1996
Earlier work this paper cites.
silence compression scheme for g. 729 optimized for terminals conforming to recommendation v. 70,
A. Itu, · 1996
Earlier work this paper cites.
Automatic segmentation, classification and clustering of broadcast news audio,
M. A. Siegler, U. Jain, B. Raj, R. M. Stern, · 1997
Earlier work this paper cites.
A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (ROVER),
J. G. Fiscus, · 1997
Earlier work this paper cites.
Partitioning and transcription of broadcast news data,
J.-L. Gauvain, L. Lamel, G. Adda, · 1998
Earlier work this paper cites.
Speaker, environment and channel change detection and clustering via the Bayesian Information Criterion,
S. S. Chen, P. S. Gopalakrishnan, · 1998
Earlier work this paper cites.
Robust detection of speech activity in the presence of noise,
R. Sarikaya, J. H. Hansen, · 1998
Earlier work this paper cites.
Speaker, environment and channel change detection and clustering via the bayesian information criterion,
S. Chen, P. Gopalakrishnan, et al., · 1998
Earlier work this paper cites.
Fast speaker change detection for broadcast news transcription and indexing,
D. Liu, F. Kubala, · 1999
Earlier work this paper cites.
Robust energy normalization using speech/nonspeech discriminator for german connected digit recognition,
R. Chengalvarayan, · 1999
Earlier work this paper cites.
A statistical model-based voice activity detection,
J. Sohn, N. S. Kim, W. Sung, · 1999
Earlier work this paper cites.
Improved speaker segmentation and segments clustering using the bayesian information criterion,
A. Tritschler, R. A. Gopinath, · 1999
Earlier work this paper cites.
Robust voice activity detection algorithm for estimating noise spectrum,
K.-H. Woo, T.-Y. Yang, K.-J. Park, C. Lee, · 2000
Earlier work this paper cites.
Strategies for automatic segmentation of audio data,
T. Kemp, M. Schmidt, M. Westphal, A. Waibel, · 2000
Earlier work this paper cites.
A speaker tracking system based on speaker turn detection for nist evaluation,
J.-F. Bonastre, P. Delacourt, C. Fredouille, T. Merlin, C. Wellekens, · 2000
Earlier work this paper cites.
Distbic: A speaker-based segmentation for audio data indexing,
P. Delacourt, C. J. Wellekens, · 2000
Earlier work this paper cites.
Speaker verification using adapted gaussian mixture models,
D. A. Reynolds, T. F. Quatieri, R. B. Dunn, · 2000
Earlier work this paper cites.
Robust voice activity detection using higher-order statistics in the lpc residual domain,
E. Nemer, R. Goubran, S. Mahmoud, · 2001
Earlier work this paper cites.
Multispeaker speech activity detection for the icsi meeting recorder,
T. Pfau, D. P. Ellis, A. Stolcke, · 2001
Earlier work this paper cites.
Speaker change detection and speaker clustering using vq distortion for broadcast news speech recognition,
K. Mori, S. Nakagawa, · 2001
Earlier work this paper cites.
On spectral clustering: Analysis and an algorithm,
A. Ng, M. Jordan, Y. Weiss, · 2001
Earlier work this paper cites.
Unsupervised speaker segmentation of telephone conversations,
A. E. Rosenberg, A. Gorin, Z. Liu, P. Parthasarathy, · 2002
Earlier work this paper cites.
Mean shift: A robust approach toward feature space analysis,
D. Comaniciu, P. Meer, · 2002
Earlier work this paper cites.
An investigation into the interactions between speaker diarisation systems and automatic speech transcription,
S. E. Tranter, K. Yu, D. A. Reynolds, G. Evermann, D. Y. Kim, P. C. Woodland, · 2003
Earlier work this paper cites.
A robust speaker clustering algorithm,
J. Ajmera, C. Wooters, · 2003
Earlier work this paper cites.
A cross-channel modeling approach for automatic segmentation of conversational telephone speech,
D. Liu, F. Kubala, · 2003
Earlier work this paper cites.
Corpus of spontaneous japanese: Its design and evaluation,
K. Maekawa, · 2003
Earlier work this paper cites.
The ICSI meeting corpus,
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, C. Wooters, · 2003
Earlier work this paper cites.
Speaker diarisation for broadcast news,
S. E. Tranter, D. A. Reynolds, · 2004
Earlier work this paper cites.
Generating and evaluating for automatic speech recognition of conversational telephone speech,
S. E. Tranter, K. Yu, G. Evermann, P. C. Woodland, · 2004
Earlier work this paper cites.
Clustering and segmenting speakers and their locations in meetings,
J. Ajmera, G. Lathoud, L. McCowan, · 2004
Earlier work this paper cites.
Speaker segmentation and clustering in meetings,
Q. Jin, K. Laskowski, T. Schultz, A. Waibel, · 2004
Earlier work this paper cites.
A novel method for two-speaker segmentation,
R. Gangadharaiah, B. Narayanaswamy, N. Balakrishnan, · 2004
Earlier work this paper cites.
Robust speaker change detection,
J. Ajmera, I. McCowan, H. Bourlard, · 2004
Earlier work this paper cites.
Speaker diarization from speech transcripts,
L. Canseco-Rodriguez, L. Lamel, J.-L. Gauvain, · 2004
Earlier work this paper cites.
The ester evaluation campaign for the rich transcription of french broadcast news.,
G. Gravier, J.-F. Bonastre, E. Geoffrois, S. Galliano, K. McTait, K. Choukri, · 2004
Earlier work this paper cites.
J. S. Garofolo, C. D. Laprun, J. G. Fiscus, The rich transcription 2004 spring meeting recognition evaluation, NIST, 2004
2004
Earlier work this paper cites.
Approaches and applications of audio diarization,
D. A. Reynolds, P. Torres-Carrasquillo, · 2005
Earlier work this paper cites.
Combining speaker identification and BIC for speaker diarization,
X. Zhu, C. Barras, S. Meignier, J.-L. Gauvain, · 2005
Earlier work this paper cites.
Eigenvoice modeling with sparse training data,
P. Kenny, G. Boulianne, P. Dumouchel, · 2005
Earlier work this paper cites.
F. Valente, Variational Bayesian methods for audio indexing, Ph.D. thesis, 2005
2005
Earlier work this paper cites.
The ami meeting corpus: A pre-announcement,
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, et al., · 2005
Earlier work this paper cites.
A generalization of blind source separation algorithms for convolutive mixtures based on second-order statistics,
H. Buchner, R. Aichner, W. Kellermann, · 2005
Earlier work this paper cites.
An overview of automatic speaker diarization systems,
S. E. Tranter, D. A. Reynolds, · 2006
Earlier work this paper cites.
Step-by-step and integrated approaches in broadcast news speaker diarization,
S. Meignier, D. Moraru, C. Fredouille, J.-F. Bonastre, L. Besacier, · 2006
Earlier work this paper cites.
Purity algorithms for speaker diarization of meetings data,
X. Anguera, C. Wooters, J. Hernando, · 2006
Earlier work this paper cites.
The rich transcription 2006 spring meeting recognition evaluation,
J. G. Fiscus, J. Ajot, M. Michel, J. S. Garofolo, · 2006
Earlier work this paper cites.
Unsupervised speaker change detection using probabilistic pattern matching,
A. Malegaonkar, A. Ariyaeeinia, P. Sivakumaran, J. Fortuna, · 2006
Earlier work this paper cites.
Fast incremental clustering of gaussian mixture speaker models for scaling up retrieval in on-line broadcast,
J. E. Rougui, M. Rziza, D. Aboutajdine, M. Gelgon, J. Martinez, · 2006
Earlier work this paper cites.
A spectral clustering approach to speaker diarization,
H. Ning, M. Liu, H. Tang, T. S. Huang, · 2006
Earlier work this paper cites.
The AMI meeting corpus: a pre-announcement,
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. P. andD. Reidsma, P. Wellner, · 2006
Earlier work this paper cites.
Speaker overlaps and ASR errors in meetings: Effects before, during, and after the overlap,
O. Cetin, E. Shriberg, · 2006
Earlier work this paper cites.
Progress in the AMIDA speaker diarization system for meeting data,
D. A. V. Leeuwen, M. Konecny, · 2007
Earlier work this paper cites.
Speaker and session variability in gmm-based speaker verification,
P. Kenny, G. Boulianne, P. Ouellet, P. Dumouchel, · 2007
Earlier work this paper cites.
A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system,
K. J. Han, S. S. Narayanan, · 2007
Earlier work this paper cites.
A tutorial on spectral clustering,
U. Von Luxburg, · 2007
Earlier work this paper cites.
The ibm rt07 evaluation systems for speaker diarization on lecture meetings,
J. Huang, E. Marcheret, K. Visweswariah, G. Potamianos, · 2007
Earlier work this paper cites.
The Rich Transcription 2007 meeting recognition evaluation,
J. Fiscus, J. Ajot, J. Garofolo, · 2007
Earlier work this paper cites.
Measuring dependence of bin-wise separated signals for permutation alignment in frequency-domain BSS,
H. Sawada, S. Araki, S. Makino, · 2007
Earlier work this paper cites.
Efficient use of overlap information in speaker diarization,
S. Otterson, M. Ostendorf, · 2007
Earlier work this paper cites.
Stream-based speaker segmentation using speaker factors and eigenvoices,
F. Castaldo, D. Colibro, E. Dalmasso, P. Laface, C. Vair, · 2008
Earlier work this paper cites.
A study of interspeaker variability in speaker verification,
P. Kenny, P. Ouellet, N. Dehak, V. Gupta, P. Dumouchel, · 2008
Earlier work this paper cites.
Bayesian analysis of speaker diarization with eigenvoice priors,
P. Kenny, · 2008
Earlier work this paper cites.
Overlapped speech detection for improved speaker diarization in multiparty meetings,
K. Boakye, B. Trueba-Hornero, O. Vinyals, G. Friedland, · 2008
Earlier work this paper cites.
An information theoretic approach to speaker diarization of meeting data,
D. Vijayasenan, F. Valente, H. Bourlard, · 2009
Earlier work this paper cites.
The majority wins: a method for combining speaker diarization systems,
M. Huijbregts, D. van Leeuwen, F. Jong, · 2009
Earlier work this paper cites.
The ester 2 evaluation campaign for the rich transcription of french radio broadcasts,
S. Galliano, G. Gravier, L. Chaubard, · 2009
Earlier work this paper cites.
Diarization of telephone conversations using factor analysis,
P. Kenny, D. Reynolds, F. Castaldo, · 2010
Earlier work this paper cites.
Variational Bayesian speaker diarization of meeting recordings,
F. Valente, P. Motlicek, D. Vijayasenan, · 2010
Earlier work this paper cites.
Speech dereverberation based on variance-normalized delayed linear prediction,
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, B.-H. Juang, · 2010
Cited alongside, same era.
An i-vector extractor suitable for speaker recognition with both microphone and telephone speech.,
M. Senoussaoui, P. Kenny, N. Dehak, P. Dumouchel, et al., · 2010
Cited alongside, same era.
Bayesian speaker verification with heavy-tailed priors.,
P. Kenny, · 2010
Cited alongside, same era.
Speaker clustering via the mean shift algorithm,
T. Stafylakis, V. Katsouros, G. Carayannis, · 2010
Cited alongside, same era.
Diarization of telephone conversations using factor analysis,
P. Kenny, D. Reynolds, F. Castaldo, · 2010
Cited alongside, same era.
System output combination for improved speaker diarization,
S. Bozonnet, N. Evans, X. Anguera, O. Vinyals, G. Friedland, C. Fredouille, · 2010
Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge.,
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, et al., · 2018
Later among the works it cites.
BUT system for DIHARD speech diarization challenge 2018.,
M. Diez, F. Landini, L. Burget, J. Rohdin, A. Silnova, K. Zmolíková, O. Novotnỳ, K. Veselỳ, O. Glembek, O. Plchot, et al., · 2018
Later among the works it cites.
Densely connected progressive learning for lstm-based speech enhancement,
T. Gao, J. Du, L.-R. Dai, C.-H. Lee, · 2018
Later among the works it cites.
NARA-WPE: A python package for weighted prediction error dereverberation in numpy and tensorflow for online and offline processing,
L. Drude, J. Heymann, C. Boeddeker, R. Haeb-Umbach, · 2018
Later among the works it cites.
Recognizing overlapped speech in meetings: A multichannel separation approach using neural networks,
T. Yoshioka, H. Erdogan, Z. Chen, X. Xiao, F. Alleva, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization,
B. Recht, M. Fazel, P. A. Parrilo, · 2010
Cited alongside, same era.
Front-end factor analysis for speaker verification,
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, P. Ouellet, · 2011
Cited alongside, same era.
Exploiting intra-conversation variability for speaker diarization,
S. Shum, N. Dehak, E. Chuangsuwanich, D. Reynolds, J. Glass, · 2011
Cited alongside, same era.
i-vector based speaker recognition on short utterances,
A. Kanagasundaram, R. Vogt, D. Dean, S. Sridharan, M. Mason, · 2011
Cited alongside, same era.
Full-covariance ubm and heavy-tailed plda in i-vector speaker verification,
P. Matějka, O. Glembek, F. Castaldo, M. J. Alam, O. Plchot, P. Kenny, L. Burget, J. Černocky, · 2011
Cited alongside, same era.
Analysis of i-vector Length Normalization in Speaker Recognition Systems,
D. Garcia-Romero, C. Y. Espy-Wilson, · 2011
Cited alongside, same era.
Front-end processing for the CHiME-5 dinner party scenario,
C. Boeddecker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, R. Haeb-Umbach, · 2018
Later among the works it cites.
Diarization is hard: some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, S. Khudanpur, · 2018
Later among the works it cites.
Relational Recurrent Neural Networks,
A. Santoro, R. Faulkner, D. Raposo, J. Rae, M. Chrzanowski, T. Weber, D. Wierstra, O. Vinyals, R. Pascanu, T. Lillicrap, · 2018
Later among the works it cites.
Single channel target speaker extraction and recognition with speaker beam,
M. Delcroix, K. Zmolikova, K. Kinoshita, A. Ogawa, T. Nakatani, · 2018
Later among the works it cites.
The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,
J. Barker, S. Watanabe, E. Vincent, J. Trmal, · 2018
Later among the works it cites.
The first dihard speech diarization challenge,
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, M. Liberman, · 2018
Later among the works it cites.
Building corpora for single-channel speech separation across multiple domains,
M. Maciejewski, G. Sell, L. P. Garcia-Perera, S. Watanabe, S. Khudanpur, · 2018
Later among the works it cites.
Meeting recognition with asynchronous distributed microphone array using block-wise refinement of mask-based MVDR beamformer,
S. Araki, N. Ono, K. Kinoshita, M. Delcroix, · 2018
Later among the works it cites.
An automated assistant for medical scribes.,
G. P. Finley, E. Edwards, A. Robinson, N. Sadoughi, J. Fone, M. Miller, D. Suendermann-Oeft, M. Brenndoerfer, N. Axtmann, · 2018
Later among the works it cites.
Fully supervised speaker diarization,
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, C. Wang, · 2019
Later among the works it cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,
Y. Luo, N. Mesgarani, · 2019
Later among the works it cites.
Enhancements for Audio-only Diarization Systems,
D. Dimitriadis, · 2019
Later among the works it cites.
Bayesian HMM based x-vector clustering for speaker diarization.,
M. Diez, L. Burget, S. Wang, J. Rohdin, J. Cernockỳ, · 2019
Later among the works it cites.
End-to-end monaural multi-speaker ASR system without pretraining,
X. Chang, Y. Qian, K. Yu, S. Watanabe, · 2019
Later among the works it cites.
Speech separation using speaker inventory,
P. Wang, Z. Chen, X. Xiao, Z. Meng, T. Yoshioka, T. Zhou, L. Lu, J. Li, · 2019
Later among the works it cites.
All-neural online source separation, counting, and diarization for meeting analysis,
T. von Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, R. Haeb-Umbach, · 2019
Later among the works it cites.
Joint Speech Recognition and Speaker Diarization via Sequence Transduction,
L. E. Shafey, H. Soltau, I. Shafran, · 2019
Later among the works it cites.
Simultaneous speech recognition and speaker diarization for monaural dialogue recordings with target-speaker acoustic models,
N. Kanda, S. Horiguchi, Y. Fujita, Y. Xue, K. Nagamatsu, S. Watanabe, · 2019
Later among the works it cites.
Speech processing for digital home assistants: Combining signal processing with deep-learning techniques,
R. Haeb-Umbach, S. Watanabe, T. Nakatani, M. Bacchiani, B. Hoffmeister, M. L. Seltzer, H. Zen, M. Souden, · 2019
Later among the works it cites.
The second DIHARD diarization challenge: Dataset, task, and baselines,
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, M. Liberman, · 2019
Later among the works it cites.
State-of-the-art speaker recognition for telephone and video speech: The JHU-MIT submission for NIST SRE18.,
J. Villalba, N. Chen, D. Snyder, D. Garcia-Romero, A. McCree, G. Sell, J. Borgstrom, F. Richardson, S. Shon, F. Grondin, et al., · 2019
Later among the works it cites.
Speaker diarization with deep speaker embeddings for dihard challenge ii.,
S. Novoselov, A. Gusev, A. Ivanov, T. Pekhovsky, A. Shulipa, A. Avdeeva, A. Gorlanov, A. Kozlov, · 2019
Later among the works it cites.
Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,
T. J. Park, K. J. Han, M. Kumar, S. Narayanan, · 2019
Later among the works it cites.
Lstm based similarity measurement with spectral clustering for speaker diarization,
Q. Lin, R. Yin, M. Li, H. Bredin, C. Barras, · 2019
Later among the works it cites.
Analysis of speaker diarization based on Bayesian HMM with eigenvoice priors,
M. Diez, L. Burget, F. Landini, J. Černockỳ, · 2019
Later among the works it cites.
DOVER: A method for combining diarization outputs,
A. Stolcke, T. Yoshioka, · 2019
Later among the works it cites.
Meeting Transcription Using Asynchronous Distant Microphones,
T. Yoshioka, D. Dimitriadis, A. Stolcke, W. Hinthorn, Z. Chen, M. Zeng, H. Xuedong, · 2019
Later among the works it cites.
Speaker diarization with lexical information,
T. J. Park, K. J. Han, J. Huang, X. He, B. Zhou, P. Georgiou, S. Narayanan, · 2019
Later among the works it cites.
End-to-end SpeakerBeam for single channel target speech recognition.,
M. Delcroix, S. Watanabe, T. Ochiai, K. Kinoshita, S. Karita, A. Ogawa, T. Nakatani, · 2019
Later among the works it cites.
VoxSRC 2019: The first VoxCeleb speaker recognition challenge,
J. S. Chung, A. Nagrani, E. Coto, W. Xie, M. McLaren, D. A. Reynolds, A. Zisserman, · 2019
Later among the works it cites.
Advances in Online Audio-Visual Meeting Transcription,
T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dimitriadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang, A. Hurvitz, L. Jiang, S. Koubi, E. Krupka, I. Leichter, C. Liu, P. Parthasarathy, A. Vinnikov, L. Wu, X. Xiao, W. Xiong, H. Wang, Z. Wang, J. Zhang, Y. Zhao, T. Zhou, · 2019
Later among the works it cites.
Guided source separation meets a strong ASR backend: Hitachi/Paderborn University joint investigation for dinner party ASR,
N. Kanda, C. Boeddeker, J. Heitkaemper, Y. Fujita, S. Horiguchi, K. Nagamatsu, R. Haeb-Umbach, · 2019
Later among the works it cites.
Speaker diarization with session-level speaker embedding refinement using graph neural networks,
J. Wang, X. Xiao, J. Wu, R. Ramamurthy, F. Rudzicz, M. Brudno, · 2020
Later among the works it cites.
Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, A. Romanenko, · 2020
Later among the works it cites.
Continuous speech separation using speaker inventory for long multi-talker recording,
C. Han, Y. Luo, C. Li, T. Zhou, K. Kinoshita, S. Watanabe, M. Delcroix, H. Erdogan, J. R. Hershey, N. Mesgarani, et al., · 2020
Later among the works it cites.
Speaker diarization with region proposal network,
Z. Huang, S. Watanabe, Y. Fujita, P. García, Y. Shao, D. Povey, S. Khudanpur, · 2020
Later among the works it cites.
Speech recognition and multi-speaker diarization of long conversations,
H. H. Mao, S. Li, J. McAuley, G. Cottrell, · 2020
Later among the works it cites.
Continuous speech separation: Dataset and analysis,
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, J. Li, · 2020
Later among the works it cites.
CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, et al., · 2020
Later among the works it cites.
The JHU multi-microphone multi-speaker asr system for the CHiME-6 challenge,
A. Arora, D. Raj, A. S. Subramanian, K. Li, B. Ben-Yair, M. Maciejewski, P. Żelasko, P. Garcia, S. Watanabe, S. Khudanpur, · 2020
Later among the works it cites.
The STC system for the CHiME-6 challenge,
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, et al., · 2020
Later among the works it cites.
Microsoft speaker diarization system for the voxceleb speaker recognition challenge 2020,
X. Xiao, N. Kanda, Z. Chen, T. Zhou, T. Yoshioka, Y. Zhao, G. Liu, J. Wu, J. Li, Y. Gong, · 2020
Later among the works it cites.
VoxSRC 2020: The second VoxCeleb speaker recognition challenge,
A. Nagrani, J. S. Chung, J. Huh, A. Brown, E. Coto, W. Xie, M. McLaren, D. A. Reynolds, A. Zisserman, · 2020
Later among the works it cites.
F. Landini, J. Profant, M. Diez, L. Burget, · 2020
Later among the works it cites.
Optimizing Bayesian HMM based x-vector clustering for the second DIHARD speech diarization challenge,
M. Diez, L. Burget, F. Landini, S. Wang, H. Černockỳ, · 2020
Later among the works it cites.
Self-attentive similarity measurement strategies in speaker diarization,
Q. Lin, Y. Hou, M. Li, · 2020
Later among the works it cites.
A memory augmented architecture for continuous speaker identification in meetings,
N. Flemotomos, D. Dimitriadis, · 2020
Later among the works it cites.
End-to-end speaker diarization as post-processing,
S. Horiguchi, P. Garcia, Y. Fujita, S. Watanabe, K. Nagamatsu, · 2020
Later among the works it cites.
Tackling real noisy reverberant meetings with all-neural source separation, counting, and diarization system,
K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, · 2020
Later among the works it cites.
End-to-end speaker diarization for an unknown number of speakers with encoder-decoder based attractors,
S. Horiguchi, Y. Fujita, S. Watanabe, Y. Xue, K. Nagamatsu, · 2020
Later among the works it cites.
Neural speaker diarization with speaker-wise chain rule,
Y. Fujita, S. Watanabe, S. Horiguchi, Y. Xue, J. Shi, K. Nagamatsu, · 2020
Later among the works it cites.
Online end-to-end neural diarization with speaker-tracing buffer,
Y. Xue, S. Horiguchi, Y. Fujita, S. Watanabe, K. Nagamatsu, · 2020
Later among the works it cites.
Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,
N. Kanda, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Zhou, T. Yoshioka, · 2020
Later among the works it cites.
Linguistically aided speaker diarization using speaker role information,
N. Flemotomos, P. Georgiou, S. Narayanan, · 2020
Later among the works it cites.
Serialized output training for end-to-end overlapped speech recognition,
N. Kanda, Y. Gaur, X. Wang, Z. Meng, T. Yoshioka, · 2020
Later among the works it cites.
Spot the conversation: Speaker diarisation in the wild,
J. S. Chung, J. Huh, A. Nagrani, T. Afouras, A. Zisserman, · 2020
Later among the works it cites.
Speaker diarization for naturalistic child-adult conversational interactions using contextual information.,
M. Kumar, S. H. Kim, C. Lord, S. Narayanan, · 2020
Later among the works it cites.
Automatic prediction of suicidal risk in military couples using multimodal interaction cues from couples conversations,
S. N. Chakravarthula, M. Nasir, S.-Y. Tseng, H. Li, T. J. Park, B. Baucom, C. J. Bryan, S. Narayanan, P. Georgiou, · 2020
Later among the works it cites.
A comprehensive evaluation of incremental speech recognition and diarization for conversational ai,
A. Addlesee, Y. Yu, A. Eshghi, · 2020
Later among the works it cites.
Third dihard challenge evaluation plan,
N. Ryant, K. Church, C. Cieri, J. Du, S. Ganapathy, M. Liberman, · 2020
Later among the works it cites.
Overlap-aware diarization: Resegmentation using neural end-to-end overlapped speech detection,
L. Bullock, H. Bredin, L. P. Garcia-Perera, · 2020
Later among the works it cites.
Exploring end-to-end multi-channel asr with bias information for meeting transcription,
X. Wang, N. Kanda, Y. Gaur, Z. Chen, Z. Meng, T. Yoshioka, · 2021
Closest in time.
Investigation of end-to-end speaker-attributed ASR for continuous multi-talker recordings,
N. Kanda, X. Chang, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Yoshioka, · 2021
Closest in time.
Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis,
D. Raj, P. Denisov, Z. Chen, H. Erdogan, Z. Huang, M. He, S. Watanabe, J. Du, T. Yoshioka, Y. Luo, N. Kanda, J. Li, S. Wisdom, J. R. Hershey, · 2021
Closest in time.
DOVER-Lap: A method for combining overlap-aware diarization outputs,
D. Raj, L. P. Garcia-Perera, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, S. Khudanpur, · 2021
Closest in time.
Multi-scale speaker diarization with neural affinity score fusion,
T. J. Park, M. Kumar, S. Narayanan, · 2021
Closest in time.
BW-EDA-EEND: Streaming end-to-end neural speaker diarization for a variable number of speakers,
E. Han, C. Lee, A. Stolcke, · 2021
Closest in time.
Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds,
K. Kinoshita, M. Delcroix, N. Tawara, · 2021
Closest in time.
Minimum Bayes risk training for end-to-end speaker-attributed ASR,
N. Kanda, Z. Meng, L. Lu, Y. Gaur, X. Wang, Z. Chen, T. Yoshioka, · 2021
Closest in time.
End-to-end speaker-attributed asr with transformer,
N. Kanda, G. Ye, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Yoshioka, · 2021
Closest in time.
Y. Fu, L. Cheng, S. Lv, Y. Jv, Y. Kong, Z. Chen, Y. Hu, L. Xie, J. Wu, H. Bu, et al., · 2021
Closest in time.
Acoustic beamforming for speaker diarization of meetings,
X. Anguera, C. Wooters, J. Hernando, · 2023
Closest in time.
Unsupervised methods for speaker diarization: An integrated and iterative approach,
S. H. Shum, N. Dehak, R. Dehak, J. R. Glass, · 2028
Closest in time.
Fusion of heterogeneous speaker recognition systems in the STBU submission for the NIST speaker recognition evaluation 2006,
N. Brummer, L. Burget, J. Cernocky, O. Glembek, F. Grezl, M. Karafiat, D. A. van Leeuwen, P. Matejka, P. Schwarz, A. Strasheim, · 2084
Closest in time.