Fetching the paper…
Reading the bibliography…
In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals.
S. S. Stevens, J. Volkmar, and E. Newman, “A scale for the measurement of the psychological magnitude pitch,”
1937
Earlier work this paper cites.
J. L. Fleiss, “Measuring nominal scale agreement among many raters,”
1971
Earlier work this paper cites.
D. Defays, “An efficient algorithm for a complete link method,”
1977
Earlier work this paper cites.
M. Slaney and M. Covell, “FaceSync: A linear operator for measuring synchronization of video facial images and audio tracks,” in
2000
Earlier work this paper cites.
W. Tong, Y. Yang, L. Jiang, S. Yu, Z. Lan, Z. Ma, W. Sze, E. Younessian, and A. G. Hauptmann, “E-lamp: Integration of innovative ideas for multimedia event detection,” ser. Machine Vision and Applications. Springer, 2002, vol. 25, pp. 5–15
2002
Earlier work this paper cites.
L. Rokach and O. Maimon, “Clustering methods,”
2005
Earlier work this paper cites.
M. Everingham, J. Sivic, and A. Zisserman, “”hello! my name is… buffy” - automatic naming of characters in tv video,” in
2006
Earlier work this paper cites.
S. Tranter and D. Reynolds, “An overview of automatic speaker diarisation systems,”
2006
Earlier work this paper cites.
M. E. Sargin, H. Aradhye, P. Moreno, and M. Zhao, “Audiovisual celebrity recognition in unconstrained web videos,” in
2009
Cited alongside, same era.
H. Jegou, M. Douze, C. Schmid, and P. Perez, “Aggregating local descriptors into a compact image representation,” in
2010
Cited alongside, same era.
“https://librivox.org/,” 2010. [Online]. Available: https://librivox.org/
2010
Cited alongside, same era.
S. Shum, N. Dehak, E. Chuangsuwanich, D. Reynolds, and J. Glass, “Exploiting intra-conversation variability for speaker diarization,” in
2011
Cited alongside, same era.
G. Monaci, “Towards real-time audiovisual speaker localization,” in
2011
Cited alongside, same era.
R. F. Lyon, “Using a cascade of asymmetric resonators with fast-acting compression as a cochlear model for machine-hearing applications,” in
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization : A review of recent research,”
2012
Later among the works it cites.
S. H. Shum, N. Dehak, R. Dehak, and J. R. Glass, “Unsupervised methods for speaker diarization: An integrated and iterative approach,”
2013
Later among the works it cites.
I. D. Gebru, S. Ba, G. Evangelidis, and R. Horaud, “Tracking the active speaker based on a joint audio-visual observation model,” in
2015
Later among the works it cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A unified embedding for face recognition and clustering,” in
2015
Later among the works it cites.
Y. Hu, J. S. Ren, J. Dai, C. Yuan, L. Xu, and W. Wang, “Deep multimodal speaker naming,” in
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
M. H. Moattar and M. M. Homayounpour, “A review on speaker diarization systems and approaches,”
2012
Cited alongside, same era.
2016
Later among the works it cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, “CNN architectures for large-scale audio classification,” 2017
2017
Closest in time.