Fetching the paper…
Reading the bibliography…
The goal of this work is to determine 'who spoke when' in real-world meetings.
J. H. DiBiase,
2000
Earlier work this paper cites.
D. Istrate, C. Fredouille, S. Meignier, L. Besacier, and J. F. Bonastre, “Nist rt’05s evaluation: pre-processing techniques and speaker diarization on multiple microphone meetings,” in
2005
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal
2005
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Acoustic beamforming for speaker diarization of meetings,”
2007
Earlier work this paper cites.
H. Hung and G. Friedland, “Towards audio-visual on-line diarization of participants in group meetings,” in
2008
Earlier work this paper cites.
G. Friedland, H. Hung, and C. Yeo, “Multi-modal speaker diarization of real-world meetings using compressed-domain video features,” in
2009
Earlier work this paper cites.
J. Schmalenstroeer, M. Kelling, V. Leutnant, and R. Haeb-Umbach, “Fusing audio and video information for online speaker diarization,” in
2009
Earlier work this paper cites.
V. Rozgic, K. J. Han, P. G. Georgiou, and S. Narayanan, “Multimodal speaker segmentation and identification in presence of overlapped speech segments,”
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
P. Matějka, O. Glembek, F. Castaldo, M. J. Alam, O. Plchot, P. Kenny, L. Burget, and J. Černocky, “Full-covariance ubm and heavy-tailed plda in i-vector speaker verification,” in
2011
Earlier work this paper cites.
G. Friedland, A. Janin, D. Imseng, X. Anguera, L. Gottlieb, M. Huijbregts, M. T. Knox, and O. Vinyals, “The icsi rt-09 speaker diarization system,”
2012
Earlier work this paper cites.
A. B. Johnston and D. C. Burnett,
2012
Cited alongside, same era.
S. Cumani, O. Plchot, and P. Laface, “Probabilistic linear discriminant analysis of i-vector posterior distributions,” in
2013
Cited alongside, same era.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in
2014
Cited alongside, same era.
Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” in
2014
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Cited alongside, same era.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” 2017
2017
Later among the works it cites.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,”
2017
Later among the works it cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” in
2017
Later among the works it cites.
2018
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,”
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
S. H. Ghalehjegh and R. C. Rose, “Deep bottleneck features for i-vector based text-independent speaker verification,” in
2015
Cited alongside, same era.
2016
Cited alongside, same era.
G. Biagetti, P. Crippa, L. Falaschetti, S. Orcioni, and C. Turchetti, “Robust speaker identification in a meeting with short audio segments,” in
2016
Cited alongside, same era.
N. Sarafianos, T. Giannakopoulos, and S. Petridis, “Audio-visual speaker diarization using fisher linear semi-discriminant analysis,”
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “SSD: Single shot multibox detector,” in
2016
Cited alongside, same era.
Later among the works it cites.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe
2018
Later among the works it cites.
P. Cabañas-Molero, M. Lucena, J. Fuertes, P. Vera-Candeas, and N. Ruiz-Reyes, “Multimodal speaker diarization for meetings using volume-evaluated srp-phat and video analysis,”
2018
Later among the works it cites.
L. Sun, J. Du, C. Jiang, X. Zhang, S. He, B. Yin, and C.-H. Lee, “Speaker diarization with enhancing speech for the first dihard challenge,”
2018
Later among the works it cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in
2018
Later among the works it cites.
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman, “VGGFace2: a dataset for recognising faces across pose and age,” in
2018
Later among the works it cites.
S.-W. Chung, J. S. Chung, and H.-G. Kang, “Perfect match: Improved cross-modal embeddings for audio-visual synchronisation,” in
2019
Closest in time.