Fetching the paper…
Reading the bibliography…
Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals.
Efficient algorithms for agglomerative hierarchical clustering methods
William HE Day and Herbert Edelsbrunner. 1984 · 1984
Earlier work this paper cites.
Relations between two sets of variates
Harold Hotelling. 1992 · 1992
Earlier work this paper cites.
Quantitative association of vocal-tract and facial behavior
Hani Yehia, Philip Rubin, and Eric Vatikiotis-Bateson. 1998 · 1998
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John Hershey and Javier Movellan. 1999 · 1999
Earlier work this paper cites.
On spectral clustering: Analysis and an algorithm. In Advances in neural information processing systems . 849–856
Andrew Y Ng, Michael I Jordan, and Yair Weiss. 2002 · 2002
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement. In International workshop on machine learning for multimodal interaction . Springer, 28–39
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al · 2005
Earlier work this paper cites.
NIST RT’05S evaluation: pre-processing techniques and speaker diarization on multiple microphone meetings. In International Workshop on Machine Learning for Multimodal Interaction . Springer, 428–439
Dan Istrate, Corinne Fredouille, Sylvain Meignier, Laurent Besacier, and Jean François Bonastre. 2005 · 2005
Earlier work this paper cites.
Front-end factor analysis for speaker verification
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet. 2010 · 2010
Earlier work this paper cites.
Bayesian speaker verification with heavy-tailed priors.. In Odyssey , Vol. 14
Patrick Kenny. 2010 · 2010
Earlier work this paper cites.
Multimodal speaker diarization
Athanasios Noulas, Gwenn Englebienne, and Ben JA Krose. 2011 · 2011
Earlier work this paper cites.
WebRTC: APIs and RTCWEB protocols of the HTML5 real-time web
Alan B Johnston and Daniel C Burnett. 2012 · 2012
Earlier work this paper cites.
Do body mass index and fat volume influence vocal quality, phonatory range, and aerodynamics in females?. In CoDAS , Vol. 25. SciELO Brasil, 310–318
Ben Barsties, Rudi Verfaillie, Nelson Roy, and Youri Maryn. 2013 · 2013
Earlier work this paper cites.
Audiovisual diarization of people in video content
Elie El Khoury, Christine Sénac, and Philippe Joly. 2014 · 2014
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition. In ICML deep learning workshop , Vol. 2. Lille
Gregory Koch, Richard Zemel, Ruslan Salakhutdinov, et al · 2015
Earlier work this paper cites.
Deep face recognition
Omkar M Parkhi, Andrea Vedaldi, and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Workshop on Multi-view Lip-reading, ACCV
J. S. Chung and A. Zisserman. 2016 · 2016
Earlier work this paper cites.
Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. In European conference on computer vision (ECCV) . Springer, 87–102
Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao. 2016 · 2016
Earlier work this paper cites.
Audio-visual speaker diarization based on spatiotemporal bayesian fusion
Israel D Gebru, Sileye Ba, Xiaofei Li, and Radu Horaud. 2017 · 2017
Earlier work this paper cites.
Learning discrete representations via information maximizing self-augmented training. In International conference on machine learning . PMLR, 1558–1567
Weihua Hu, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, and Masashi Sugiyama. 2017 · 2017
Earlier work this paper cites.
End-to-end face detection and cast grouping in movies using erdos-renyi clustering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 5276–5285
SouYoung Jin, Hang Su, Chris Stauffer, and Erik Learned-Miller. 2017 · 2017
Earlier work this paper cites.
Multimodal speaker clustering in full length movies
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis, Geoffroy Peeters, Laurent Benaroya, and Ioannis Pitas. 2017 · 2017
Earlier work this paper cites.
VoxCeleb: a large-scale speaker identification dataset. In INTERSPEECH
A. Nagrani, J. S. Chung, and A. Zisserman. 2017 · 2017
Cited alongside, same era.
S3fd: Single shot scale-invariant face detector. In Proceedings of the IEEE international conference on computer vision (ICCV) . 192–201
Shifeng Zhang, Xiangyu Zhu, Zhen Lei, Hailin Shi, Xiaobo Wang, and Stan Z Li. 2017 · 2017
Cited alongside, same era.
Deep Lip Reading: a comparison of models and an online application. In INTERSPEECH
T. Afouras, J. S. Chung, and A. Zisserman. 2018 · 2018
Cited alongside, same era.
Vggface2: A dataset for recognising faces across pose and age. In 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018) . IEEE, 67–74
Qiong Cao, Li Shen, Weidi Xie, Omkar M Parkhi, and Andrew Zisserman. 2018 · 2018
Cited alongside, same era.
VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH
J. S. Chung, A. Nagrani, and A. Zisserman. 2018 · 2018
Cited alongside, same era.
Optimizing bayesian hmm based x-vector clustering for the second dihard speech diarization challenge. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6519–6523
Mireia Diez, Lukáš Burget, Federico Landini, Shuai Wang, and Honza Černockỳ. 2020 · 2020
Later among the works it cites.
Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID. In Advances in Neural Information Processing Systems
Yixiao Ge, Feng Zhu, Dapeng Chen, Rui Zhao, and Hongsheng Li. 2020 · 2020
Later among the works it cites.
End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors. In Proc. Interspeech 2020 . 269–273
Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue, and Kenji Nagamatsu. 2020 · 2020
Later among the works it cites.
Speaker diarization with region proposal network. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6514–6518
Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola García, Yiwen Shao, Daniel Povey, and Sanjeev Khudanpur. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge.. In INTERSPEECH . 2808–2812
Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesús Villalba, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe, et al · 2018
Cited alongside, same era.
X-vectors: Robust dnn embeddings for speaker recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5329–5333
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. 2018 · 2018
Cited alongside, same era.
Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 1199–1208
Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. 2018 · 2018
Cited alongside, same era.
Generalized end-to-end loss for speaker verification. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4879–4883
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno. 2018 · 2018
Cited alongside, same era.
Speaker diarization with LSTM. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5239–5243
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno. 2018 · 2018
Cited alongside, same era.
Disjoint Mapping Network for Cross-modal Matching of Voices and Faces. In International Conference on Learning Representations
Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj, and Rita Singh. 2018 · 2018
Cited alongside, same era.
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3965–3969
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang. 2019a · 2019
Cited alongside, same era.
Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario. In INTERSPEECH
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Khokhlov, Mariya Korenevskaya, Ivan Sorokin, Tatiana Timofeeva, Anton Mitrofanov, Andrei Andrusenko, Ivan Podluzhny, et al · 2020
Later among the works it cites.
Ava active speaker: An audio-visual dataset for active speaker detection. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4492–4496
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al · 2020
Later among the works it cites.
Online multi-modal person search in videos. In European Conference on Computer Vision (ECCV) . Springer, 174–190
Jiangyue Xia, Anyi Rao, Qingqiu Huang, Linning Xu, Jiangtao Wen, and Dahua Lin. 2020 · 2020
Later among the works it cites.
Social adaptive module for weakly-supervised group activity recognition. In European Conference on Computer Vision . Springer, 208–224
Rui Yan, Lingxi Xie, Jinhui Tang, Xiangbo Shu, and Qi Tian. 2020 · 2020
Later among the works it cites.
APES: Audiovisual Person Search in Untrimmed Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 1720–1729
Juan Leon Alcazar, Fabian Caba, Long Mai, Federico Perazzi, Joon-Young Lee, Pablo Arbelaez, and Bernard Ghanem. 2021 · 2021
Closest in time.
Face, Body, Voice: Video Person-Clustering with Multiple Modalities
Andrew Brown, Vicky Kalogeiton, and Andrew Zisserman. 2021 · 2021
Closest in time.
VisualVoice: Audio-Visual Speech Separation With Cross-Modal Consistency. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 15495–15505
Ruohan Gao and Kristen Grauman. 2021 · 2021
Closest in time.
Analysis of the but diarization system for voxconverse challenge. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5819–5823
Federico Landini, Ondřej Glembek, Pavel Matějka, Johan Rohdin, Lukáš Burget, Mireia Diez, and Anna Silnova. 2021 · 2021
Closest in time.
Audio-Visual Deep Neural Network for Robust Person Verification
Yanmin Qian, Zhengyang Chen, and Shuai Wang. 2021 · 2021
Closest in time.
A Multi-View Approach to Audio-Visual Speaker Verification. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6194–6198
Leda Sarı, Kritika Singh, Jiatong Zhou, Lorenzo Torresani, Nayan Singhal, and Yatharth Saraf. 2021 · 2021
Closest in time.
Sparse r-cnn: End-to-end object detection with learnable proposals. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 14454–14463
Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al · 2021
Closest in time.
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection. In Proceedings of the 29th ACM International Conference on Multimedia . 3927–3935
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li. 2021 · 2021
Closest in time.
Seeking the shape of sound: An adaptive framework for learning voice-face association. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 16347–16356
Peisong Wen, Qianqian Xu, Yangbangyan Jiang, Zhiyong Yang, Yuan He, and Qingming Huang. 2021 · 2021
Closest in time.
Microsoft speaker diarization system for the VoxCeleb speaker recognition challenge 2020. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5824–5828
Xiong Xiao, Naoyuki Kanda, Zhuo Chen, Tianyan Zhou, Takuya Yoshioka, Sanyuan Chen, Yong Zhao, Gang Liu, Yu Wu, Jian Wu, et al · 2021
Closest in time.
Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18995–19012
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Closest in time.
Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks
Federico Landini, Ján Profant, Mireia Diez, and Lukáš Burget. 2022 · 2022
Closest in time.