Fetching the paper…
Reading the bibliography…
Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction.
Measuring nominal scale agreement among many raters
J. L. Fleiss · 1971
Earlier work this paper cites.
Look who’s talking: Speaker detection using video and audio correlation
R. Cutler and L. Davis · 2000
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
J. W. Fisher III, T. Darrell, W. T. Freeman, and P. A. Viola · 2001
Earlier work this paper cites.
Facesync: A linear operator for measuring synchronization of video facial images and audio tracks
M. Slaney and M. Covell · 2001
Earlier work this paper cites.
CUAVE: A new audio-visual database for multimodal human-computer interface research
E. K. Patterson, S. Gurbuz, Z. Tufekci, and J. N. Gowdy · 2002
Earlier work this paper cites.
Speaker localisation using audio-visual synchrony: An empirical study
H. J. Nock, G. Iyengar, and C. Neti · 2003
Earlier work this paper cites.
A segment-based audio-visual speech recognizer: Data collection, development, and initial experiments
T. J. Hazen, K. Saenko, C.-H. La, and J. R. Glass · 2004
Earlier work this paper cites.
Cross-modal analysis of audio-visual programs for speaker detection
D. Li, C. M. Taskiran, N. Dimitrova, W. Wang, M. Li, and I. K. Sethi · 2005
Earlier work this paper cites.
The AMI meeting corpus
I. McCowan, J. Carletta, W. Kraaij, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, W. Post, D. Reidsma, and P. Wellner · 2005
Earlier work this paper cites.
Visual speech recognition with loosely synchronized feature streams
K. Saenko, K. Livescu, M. Siracusa, K. Wilson, J. Glass, and T. Darrell · 2005
Earlier work this paper cites.
Robust speaker diarization for meetings: ICSI RT06s evaluation system
X. Anguera, C. Wooters, and J. Hernando · 2006
Earlier work this paper cites.
“Hello! my name is… buffy” – automatic naming of characters in TV video
M. Everingham, J. Sivic, and A. Zisserman · 2006
Earlier work this paper cites.
Audio segmentation and speaker localization in meeting videos
H. Vajaria, T. Islam, S. Sarkar, R. Sankar, and R. Kasturi · 2006
Earlier work this paper cites.
Movie/script: Alignment and parsing of video and text transcription
T. Cour, C. Jordan, E. Miltsakaki, and B. Taskar · 2008
Earlier work this paper cites.
Exploring co-occurence between speech and body movement for audio-guided video localization
H. Vajaria, S. Sarkar, and R. Kasturi · 2008
Earlier work this paper cites.
Boosting-based multimodal speaker detection for distributed meeting videos
C. Zhang, P. Yin, Y. Rui, R. Cutler, P. Viola, X. Sun, N. Pinto, and Z. Zhang · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Taking the bite out of automatic naming of characters in TV video
M. Everingham, J. Sivic, and A. Zisserman · 2009
Cited alongside, same era.
Visual speaker localization aided by acoustic models
G. Friedland, C. Yeo, and H. Hung · 2009
Cited alongside, same era.
The ESTER-2 evaluation campaign for the rich transcription of french radio broadcasts
S. Galliano, G. Gravier, and L. Chaubard · 2009
Cited alongside, same era.
Speech/non-speech detection in meetings from automatically extracted low resolution visual features
H. Hung and S. O. Ba · 2009
Cited alongside, same era.
Speaker adaptation techniques for automatic speech recognition
K. Shinoda · 2011
Cited alongside, same era.
The REPERE corpus : a multimodal corpus for person recognition
A. Giraudel, M. Carre, V. Mapelli, J. Kahn, O. Galibert, and L. Quintard · 2012
Look who’s talking: Visual identification of the active speaker in multi-party human-robot interaction
K. Stefanov, A. Sugimoto, and J. Beskow · 2016
Later among the works it cites.
Audio-visual speaker diarization based on spatiotemporal bayesian fusion
I. D. Gebru, S. Ba, X. Li, and R. Horaud · 2017
Later among the works it cites.
Automated curriculum learning for neural networks
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Later among the works it cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Later among the works it cites.
Vision-based active speaker detection in multiparty interaction
K. Stefanov, J. Beskow, and G. Salvi · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Speaker diarization of broadcast news in albayzin 2010 evaluation campaign
M. Zelenak, H. Schulz, and J. Hernando · 2012
Cited alongside, same era.
Semi-supervised learning with constraints for person identification in multimedia data
M. Bäuml, M. Tapaswi, and R. Stiefelhagen · 2013
Cited alongside, same era.
Who’s speaking? audio-supervised classification of active speakers in video
P. Chakravarty, S. Mirzaei, T. Tuytelaars, and H. Vanhamme · 2015
Cited alongside, same era.
Deep multimodal speaker naming
Y. Hu, J. S. Ren, J. Dai, C. Yuan, L. Xu, and W. Wang · 2015
Cited alongside, same era.
A convolutional neural network cascade for face detection
H. Li, Z. Lin, X. Shen, J. Brandt, and G. Hua · 2015
Cited alongside, same era.
Cross-modal supervision for learning active speaker detection in video
P. Chakravarty and T. Tuytelaars · 2016
Cited alongside, same era.
Bimodal recurrent neural network for audiovisual voice activity detection
F. Tao and C. Busso · 2017
Later among the works it cites.
The conversation: Deep audio-visual speech enhancement
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Later among the works it cites.
AVA-Speech: A densely labeled dataset of speech activity in movies
S. Chaudhuri, J. Roth, D. Ellis, A. C. Gallagher, L. Kaver, R. Marvin, C. Pantofaru, N. C. Reale, L. G. Reid, K. Wilson, and Z. Xi · 2018
Later among the works it cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Later among the works it cites.
Visual speech enhancement
A. Gabbay, A. Shamir, and S. Peleg · 2018
Later among the works it cites.
Audio-visual speaker diarization based on spatiotemporal bayesian fusion
I. D. Gebru, S. Ba, X. Li, and R. Horaud · 2018
Later among the works it cites.
AVA: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2018
Later among the works it cites.
Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers
K. Hoover, S. Chaudhuri, C. Pantofaru, M. Slaney, and I. Sturdy · 2018
Later among the works it cites.
Voxceleb: a large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Later among the works it cites.
Large-scale visual speech recognition
B. Shillingford, Y. Assael, M. W. Hoffman, T. Paine, C. Hughes, U. Prabhu, H. Liao, H. Sak, K. Rao, L. Bennett, M. Mulville, B. Coppin, B. Laurie, A. Senior, and N. de Freitas · 2018
Later among the works it cites.