Fetching the paper…
Reading the bibliography…
Active speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates.
R. Cutler and L. Davis, “Look who’s talking: Speaker detection using video and audio correlation,” in IEEE International Conference on Multimedia and Expo , vol. 3, 2000, pp. 1589–1592
2000
Earlier work this paper cites.
T. Strybel and K. Fujimoto, “Minimum audible angles in the horizontal and vertical planes: Effects of stimulus onset asynchrony and burst duration.” The Journal of the Acoustical Society of America , vol. 108 6, pp. 3092–5, 2000
2000
Earlier work this paper cites.
X. Zhu, A. B. Goldberg, R. Brachman, and T. Dietterich, Introduction to Semi-Supervised Learning . Morgan and Claypool Publishers, 2009
2009
Earlier work this paper cites.
F. Haider and S. A. Moubayed, “Towards speaker detection using lips movements for human-machine multiparty dialogue,” in Proceedings of Fonetik , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems , vol. 25, no. 2, pp. 84–90, 2012
2012
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NIPS Deep Learning and Representation Learning Workshop , 2015
2015
Earlier work this paper cites.
P. Chakravarty, S. Mirzaei, T. Tuytelaars, and H. V. hamme, “Who’s speaking?: Audio-supervised classification of active speakers in video,” ACM International Conference on Multimodal Interaction , 2015
2015
Earlier work this paper cites.
M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International Journal of Computer Vision , vol. 111, no. 1, pp. 98–136, 2015
2015
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in NIPS , 2016
2016
Earlier work this paper cites.
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in ECCV , 2016
2016
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in ICCV , 2017, pp. 609–617
2017
Earlier work this paper cites.
S. Wu, M. Kan, Z. He, S. Shan, and X. Chen, “Funnel-structured cascade for multi-view face detection with alignment-awareness,” Neurocomputing , 2017
2017
Cited alongside, same era.
T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in INTERSPEECH , 2018
2018
Cited alongside, same era.
A. Ephrat et al
2018
Cited alongside, same era.
T. L. Pedro Morgado, Nuno Vasconcelos and O. Wang, “Self-supervised generation of spatial audio for 360° video,” in NIPS , 2018
2018
Cited alongside, same era.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,” in ECCV , 2018
2018
Cited alongside, same era.
Y. LeCun. (2019, Apr) Available at: ” https://twitter.com/ylecun/status/1123235709802905600 ” and at: ” https://www.facebook.com/722677142/posts/10155934004262143/ ”
2019
Later among the works it cites.
2019
Later among the works it cites.
A. B. Vasudevan, D. Dai, and L. V. Gool, “Semantic object prediction and spatial sound super-resolution with binaural sounds,” in ECCV , 2020, pp. 638–655
2020
Later among the works it cites.
A. Politis, S. Adavanne, and T. Virtanen, “A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2020
2020
Later among the works it cites.
J. Roth et al
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Stenzel and P. J. B. Jackson, “Perceptual thresholds of audio-visual spatial coherence for a variety of audio-visual objects,” in Conference on Audio for Virtual and Augmented Reality , 2018
2018
Cited alongside, same era.
R. Gao and K. Grauman, “2.5D visual sound,” in CVPR , 2019
2019
Cited alongside, same era.
R. Gao and K. Grauman, “Co-separating sounds of visual objects,” in ICCV , 2019
2019
Cited alongside, same era.
C. Gan, H. Zhao, P. Chen, D. Cox, and A. Torralba, “Self-supervised moving vehicle tracking with stereo sound,” ICCV , pp. 7052–7061, 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
J. L. Alcazar et al
2020
Later among the works it cites.
M. Blanco Galindo, P. Coleman, and P. J. B. Jackson, “Microphone array geometries for horizontal spatial audio object capture with beamforming,” Journal of the Audio Engineering Society , vol. 68, no. 5, pp. 324–337, 2020
2020
Later among the works it cites.
D. Berghi, H. Stenzel, M. Volino, A. Hilton, and P. J. B. Jackson, “Audio-visual spatial alignment requirements of central and peripheral object events,” in 2020 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW) , 2020, pp. 666–667
2020
Later among the works it cites.
F. Rivera Valverde, J. Valeria Hurtado, and A. Valada, “There is more than meets the eye: Self-supervised multi-object detection and tracking with sound by distilling multimodal knowledge,” in CVPR , 2021
2021
Later among the works it cites.
M. A. Mohd Izhar, M. Volino, A. Hilton, and P. J. B. Jackson, “Tracking sound sources for object-based spatial audio in 3D audio-visual production,” in Forum Acusticum , 2020, pp. 2051–2058
2058
Closest in time.