Fetching the paper…
Reading the bibliography…
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers.
Bimodal recurrent neural network for audiovisual voice activity detection.. In INTERSPEECH . 1938–1942
Fei Tao and Carlos Busso. 2017 · 1942
Earlier work this paper cites.
Speech rates in british english
Steve Tauroza and Desmond Allison. 1990 · 1990
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping. In European conference on computer vision (ECCV) . Springer, 25–36
Thomas Brox, Andrés Bruhn, Nils Papenberg, and Joachim Weickert. 2004 · 2004
Earlier work this paper cites.
A duality based approach for realtime tv-l 1 optical flow. In Joint pattern recognition symposium . Springer, 214–223
Christopher Zach, Thomas Pock, and Horst Bischof. 2007 · 2007
Earlier work this paper cites.
The Kaldi speech recognition toolkit. In IEEE 2011 Workshop on Automatic Speech Recognition and Understanding (ASRU)
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Who’s speaking? Audio-supervised classification of active speakers in video. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction (ICMI ) . 87–90
Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, and Hugo Van hamme. 2015 · 2015
Earlier work this paper cites.
Multimodal multi-channel on-line speaker diarization using sensor fusion through SVM
Vicente Peruffo Minotto, Claudio Rosito Jung, and Bowon Lee. 2015 · 2015
Earlier work this paper cites.
MUSAN: A Music, Speech, and Noise Corpus
D. Snyder, G. Chen, and D. Povey. 2015 · 2015
Earlier work this paper cites.
Cross-modal supervision for learning active speaker detection in video. In European Conference on Computer Vision (ECCV) . Springer, 285–301
Punarjay Chakravarty and Tinne Tuytelaars. 2016 · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Asian conference on computer vision (ACCV) . Springer, 251–263
Joon Son Chung and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
Visual voice activity detection in the wild
Foteini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikolaos Nikolaidis, and Ioannis Pitas. 2016 · 2016
Earlier work this paper cites.
A study on data augmentation of reverberant speech for robust speech recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2017 . 5220–5224
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur. 2017 · 2017
Earlier work this paper cites.
S3FD: Single shot scale-invariant face detector. In Proceedings of the IEEE international conference on computer vision (ICCV) . 192–201
Shifeng Zhang, Xiangyu Zhu, Zhen Lei, Hailin Shi, Xiaobo Wang, and Stan Z Li. 2017 · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2018c · 2018
Earlier work this paper cites.
The Conversation: deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018a · 2018
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018b · 2018
Earlier work this paper cites.
VoxCeleb2: Deep speaker recognition. In Proc. of Interspeech 2018 . 1086–1090
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T. Freeman, and Michael Rubinstein. 2018 · 2018
Cited alongside, same era.
Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 7132–7141
Jie Hu, Li Shen, and Gang Sun. 2018 · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European Conference on Computer Vision (ECCV) . 631–648
Andrew Owens and Alexei A Efros. 2018 · 2018
Cited alongside, same era.
A convolutional neural network smartphone app for real-time voice activity detection
Abhishek Sehgal and Nasser Kehtarnavaz. 2018 · 2018
Cited alongside, same era.
X-Vectors: robust DNN embeddings for speaker recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2018 . 5329–5333
RealVAD: A real-world dataset and a method for voice activity detection by body motion analysis
Cigdem Beyan, Muhammad Shahid, and Vittorio Murino. 2020 · 2020
Later among the works it cites.
Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning. In Proceedings of the 28th ACM International Conference on Multimedia (ACM MM) . 3884–3892
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, and Yuejie Zhang. 2020 · 2020
Later among the works it cites.
In defence of metric learning for speaker recognition. In Interspeech 2020
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, and Icksang Han. 2020a · 2020
Later among the works it cites.
Spot the Conversation: Speaker Diarisation in the Wild. In Proc. Interspeech 2020 . 299–303
Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, and Andrew Zisserman. 2020b · 2020
Later among the works it cites.
Oxford guide to plain English
Martin Cutts. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur. 2018 · 2018
Cited alongside, same era.
An end-to-end multimodal voice activity detection using wavenet encoder and residual networks
Ido Ariav and Israel Cohen. 2019 · 2019
Cited alongside, same era.
Naver at ActivityNet Challenge 2019–Task B Active Speaker Detection (AVA)
Joon Son Chung. 2019 · 2019
Cited alongside, same era.
Who Said That? Audio-Visual Speaker Diarisation of Real-World Meetings. In Proc. Interspeech 2019 . 371–375
Joon Son Chung, Bong-Jin Lee, and Icksang Han. 2019b · 2019
Cited alongside, same era.
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2019 . IEEE, 3965–3969
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang. 2019a · 2019
Cited alongside, same era.
Randomly weighted cnns for (music) audio classification. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2019 . IEEE, 336–340
Jordi Pons and Xavier Serra. 2019 · 2019
Cited alongside, same era.
Comparisons of visual activity primitives for voice activity detection. In International Conference on Image Analysis and Processing . Springer, 48–59
Muhammad Shahid, Cigdem Beyan, and Vittorio Murino. 2019 · 2019
Cited alongside, same era.
Leveraging long-range temporal relationships between proposals for video object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 9756–9764
Mykhailo Shvets, Wei Liu, and Alexander C Berg. 2019 · 2019
Cited alongside, same era.
Personal VAD: Speaker-Conditioned Voice Activity Detection. In Proc. Odyssey 2020 The Speaker and Language Recognition Workshop . 433–439
Shaojin Ding, Quan Wang, Shuo-Yiin Chang, Li Wan, and Ignacio Lopez Moreno. 2020 · 2020
Later among the works it cites.
Improved active speaker detection based on optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops . 950–951
Chong Huang and Kazuhito Koishida. 2020 · 2020
Later among the works it cites.
Muse: Multi-modal target speaker extraction with visual cues
Zexu Pan, Ruijie Tao, Chenglin Xu, and Haizhou Li. 2020 · 2020
Later among the works it cites.
AVA active speaker: An audio-visual dataset for active speaker detection. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2020 . IEEE, 4492–4496
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al · 2020
Later among the works it cites.
End to end lip synchronization with a temporal autoencoder. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 341–350
Yoav Shalev and Lior Wolf. 2020 · 2020
Later among the works it cites.
Crossmodal learning for audio-visual speech event localization
Rahul Sharma, Krishna Somandepalli, and Shrikanth Narayanan. 2020 · 2020
Later among the works it cites.
Detecting Aedes aegypti mosquitoes through audio classification with convolutional neural networks
Marcelo Schreiber Fernandes, Weverton Cordeiro, and Mariana Recamonde-Mendoza. 2021 · 2021
Closest in time.
MAAS: Multi-modal assignation for active speaker detection
Juan León-Alcázar, Fabian Caba Heilbron, Ali Thabet, and Bernard Ghanem. 2021 · 2021
Closest in time.
Audio-visual tracking of concurrent speakers
Xinyuan Qian, Alessio Brutti, Oswald Lanz, Maurizio Omologo, and Andrea Cavallaro. 2021a · 2021
Closest in time.
Multi-target DoA estimation with an audio-visual fusion mechanism. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2021 . IEEE, 4280–4284
Xinyuan Qian, Maulik Madhavi, Zexu Pan, Jiadong Wang, and Haizhou Li. 2021b · 2021
Closest in time.
S-VVAD: Visual Voice Activity Detection by Motion Segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 2332–2341
Muhammad Shahid, Cigdem Beyan, and Vittorio Murino. 2021 · 2021
Closest in time.