Fetching the paper…
Reading the bibliography…
The use of audio and visual modality for speaker localization has been well studied in the literature by exploiting their complementary characteristics.
C. Knapp and G. Carter, “The generalized correlation method for estimation of time delay,” IEEE transactions on acoustics, speech, and signal processing
1976
Earlier work this paper cites.
M. S. Brandstein and H. F. Silverman, “A robust method for speech signal time-delay estimation in reverberant rooms,” in 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing
1997
Earlier work this paper cites.
Y. Rubner, C. Tomasi, and L. J. Guibas, “The earth mover’s distance as a metric for image retrieval,” International journal of computer vision
2000
Earlier work this paper cites.
D. Zotkin, R. Duraiswami, and L. S. Davis, “Multimodal 3-d tracking and event detection via the particle filter,” in Proceedings IEEE workshop on detection and recognition of events in video
2001
Earlier work this paper cites.
G. Lathoud, J.-M. Odobez, and D. Gatica-Perez, “Av16. 3: An audio-visual corpus for speaker localization and tracking,” in International Workshop on Machine Learning for Multimodal Interaction
2004
Earlier work this paper cites.
A. Hampapur, L. Brown, J. Connell, A. Ekin, N. Haas, M. Lu, H. Merkl, and S. Pankanti, “Smart video surveillance: exploring the concept of multiscale spatiotemporal tracking,” IEEE signal processing magazine
2005
Earlier work this paper cites.
S. T. Birchfield and S. Rangarajan, “Spatiograms versus histograms for region-based tracking,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05)
2005
Earlier work this paper cites.
CRC press, 2007
P. C. Loizou, Speech enhancement: theory and practice · 2007
Earlier work this paper cites.
A. Brutti, M. Omologo, and P. Svaizer, “Localization of multiple speakers based on a two step acoustic map analysis,” in 2008 IEEE International Conference on Acoustics, Speech and Signal Processing
2008
Earlier work this paper cites.
E. D’Arca, N. M. Robertson, and J. Hopgood, “Person tracking via audio and video fusion,” 2012
2012
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “Demand: a collection of multi-channel recordings of acoustic noise in diverse environments,” in Proc. Meetings Acoust
2013
Earlier work this paper cites.
T.-H. Vu, A. Osokin, and I. Laptev, “Context-aware cnns for person head detection,” in Proceedings of the IEEE International Conference on Computer Vision
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP)
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2016
Earlier work this paper cites.
R. Essaadali and A. Kouki, “A new simple unmanned aerial vehicle doppler effect rf reducing technique,” in MILCOM 2016-2016 IEEE Military Communications Conference
2016
Earlier work this paper cites.
I. D. Gebru, X. Alameda-Pineda, F. Forbes, and R. Horaud, “Em algorithms for weighted-data clustering with application to audio-visual scene analysis,” IEEE transactions on pattern analysis and machine intelligence
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” arXiv preprint arXiv:1804.02767
2018
Cited alongside, same era.
W. He, P. Motlicek, and J.-M. Odobez, “Deep neural networks for multiple speaker detection and localization,” in 2018 IEEE International Conference on Robotics and Automation (ICRA)
2018
Cited alongside, same era.
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al
2018
Cited alongside, same era.
J. M. Vera-Diaz, D. Pizarro, and J. Macias-Guarasa, “Towards end-to-end acoustic localization using deep learning: From audio signals to source position coordinates,” Sensors
2018
Cited alongside, same era.
S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen, “Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,” IEEE Journal of Selected Topics in Signal Processing
2021
Later among the works it cites.
D. Michelsanti, Z.-H. Tan, S.-X. Zhang, Y. Xu, M. Yu, D. Yu, and J. Jensen, “An overview of deep-learning-based audio-visual speech enhancement and separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing
2021
Later among the works it cites.
X. Qian, M. Madhavi, Z. Pan, J. Wang, and H. Li, “Multi-target doa estimation with an audio-visual fusion mechanism,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2021
Later among the works it cites.
J. Wissing, B. Boenninghoff, D. Kolossa, T. Ochiai, M. Delcroix, K. Kinoshita, T. Nakatani, S. Araki, and C. Schymura, “Data fusion for audiovisual speaker localization: Extending dynamic stream weights to the spatial domain,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
J. Li, Y. Wang, C. Wang, Y. Tai, J. Qian, J. Yang, C. Wang, J. Li, and F. Huang, “Dsfd: dual shot face detector,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2019
Cited alongside, same era.
X. Qian, A. Brutti, O. Lanz, M. Omologo, and A. Cavallaro, “Multi-speaker tracking from an audio–visual sensing device,” IEEE Transactions on Multimedia
2019
Cited alongside, same era.
2019
Cited alongside, same era.
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Epic-fusion: Audio-visual temporal binding for egocentric action recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision
2019
Cited alongside, same era.
Y. Wu, L. Zhu, Y. Yan, and Y. Yang, “Dual attention matching for audio-visual event localization,” in Proceedings of the IEEE/CVF international conference on computer vision
2019
Cited alongside, same era.
M. Liu, S. Tang, Y. Li, and J. M. Rehg, “Forecasting human-object interaction: Joint prediction of motor attention and egocentric activity,” Computer Vision–ECCV 2020
2020
Cited alongside, same era.
C. Northcutt, S. Zha, S. Lovegrove, and R. Newcombe, “Egocom: A multi-person multi-modal egocentric communications dataset,” IEEE Transactions on Pattern Analysis and Machine Intelligence
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Berghi, A. Hilton, and P. J. Jackson, “Visually supervised speaker detection and localization via microphone array,” in 2021 IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP)
2021
Later among the works it cites.
J. Ong, B. T. Vo, S. Nordholm, B.-N. Vo, D. Moratuwage, and C. Shim, “Audio-visual based online multi-source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing
2022
Later among the works it cites.
E. Z. Xu, Z. Song, S. Tsutsui, C. Feng, M. Ye, and M. Z. Shou, “Ava-avd: Audio-visual speaker diarization in the wild,” in Proceedings of the 30th ACM International Conference on Multimedia
2022
Later among the works it cites.
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al
2022
Later among the works it cites.
H. Jiang, C. Murdock, and V. K. Ithapu, “Egocentric deep multi-channel audio-visual active speaker localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2022
Later among the works it cites.
P.-A. Grumiaux, S. Kitić, L. Girin, and A. Guérin, “A survey of sound source localization with deep learning methods,” The Journal of the Acoustical Society of America
2022
Later among the works it cites.
X. Qian, Q. Zhang, G. Guan, and W. Xue, “Deep audio-visual beamforming for speaker localization,” IEEE Signal Processing Letters
2022
Later among the works it cites.
X. Qian, Z. Wang, J. Wang, G. Guan, and H. Li, “Audio-visual cross-attention network for robotic speaker tracking,” IEEE/ACM Transactions on Audio, Speech, and Language Processing
2022
Later among the works it cites.
A. S. Subramanian, C. Weng, S. Watanabe, M. Yu, and D. Yu, “Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognition,” Computer Speech & Language
2022
Later among the works it cites.
Y. Li, H. Liu, and H. Tang, “Multi-modal perception attention network with self-supervised learning for audio-visual speaker tracking,” in Proceedings of the AAAI Conference on Artificial Intelligence
2022
Later among the works it cites.
Y. Wu, R. Hu, X. Wang, and S. Ke, “Multi-speaker doa estimation using audio and visual modality,” Neural Processing Letters
2023
Closest in time.