Fetching the paper…
Reading the bibliography…
An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is talking, and not.
J. Hershey and J. Movellan, “Audio vision: Using audio-visual synchrony to locate sounds,” in Advances in Neural Information Processing Systems , S. Solla, T. Leen, and K. Müller, Eds., vol. 12. MIT Press, 2000. [Online]. Available: https://proceedings.neurips.cc/paper/1999/file/b618c3210e934362ac261db280128c22-Paper.pdf
1999
Earlier work this paper cites.
J. W. Fisher III, T. Darrell, W. T. Freeman, and P. A. Viola, “Learning joint statistical models for audio-visual fusion and segregation,” in Advances in neural information processing systems , 2001, pp. 772–778
2001
Earlier work this paper cites.
M. Everingham, J. Sivic, and A. Zisserman, “Hello! my name is… buffy”–automatic naming of characters in tv video.” in BMVC , vol. 2, no. 4, 2006, p. 6
2006
Earlier work this paper cites.
Z. Barzelay and Y. Y. Schechner, “Harmony in motion,” in 2007 IEEE Conference on Computer Vision and Pattern Recognition , 2007, pp. 1–8
2007
Earlier work this paper cites.
L. Shams and R. Kim, “Crossmodal influences on visual perception,” Physics of life reviews , pp. 269–284, 2010
2010
Earlier work this paper cites.
J. Klemen and C. D. Chambers, “Current perspectives and methods in studying neural mechanisms of multisensory interactions,” Neuroscience & Biobehavioral Reviews , pp. 111–133, 2012
2012
Earlier work this paper cites.
K. Schmiedchen, C. Freigang, I. Nitsche, and R. Rübsamen, “Crossmodal interactions and multisensory integration in the perception of audio-visual motion—a free-field study,” Brain research , pp. 99–111, 2012
2012
Earlier work this paper cites.
M. Hermans and B. Schrauwen, “Training and analysing deep recurrent neural networks,” in Advances in neural information processing systems , 2013, pp. 190–198
2013
Earlier work this paper cites.
P. Chakravarty, S. Mirzaei, T. Tuytelaars, and H. Van hamme, “Who’s speaking? audio-supervised classification of active speakers in video,” in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction , 2015, pp. 87–90
2015
Earlier work this paper cites.
L. Bazzani, A. Bergamo, D. Anguelov, and L. Torresani, “Self-taught object localization with deep networks,” in 2016 IEEE winter conference on applications of computer vision (WACV) . IEEE, 2016, pp. 1–9
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Earlier work this paper cites.
H. Bilen and A. Vedaldi, “Weakly supervised deep detection networks,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 2846–2854
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 251–263
2016
Earlier work this paper cites.
P. Chakravarty, J. Zegers, T. Tuytelaars, and H. Van hamme, “Active speaker detection with audio-visual co-training,” in Proceedings of the 18th ACM International Conference on Multimodal Interaction , 2016, pp. 312–316
2016
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 609–617
2017
Cited alongside, same era.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in 2017 IEEE International Conference on Computer Vision (ICCV) , Oct 2017, pp. 618–626
2017
Cited alongside, same era.
P. Tang, X. Wang, X. Bai, and W. Liu, “Multiple instance detection network with online instance classifier refinement,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 2843–2851
2017
Cited alongside, same era.
X. Wang, A. Shrivastava, and A. Gupta, “A-fast-rcnn: Hard positive generation via adversary for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2606–2615
2017
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, “The sound of motions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 1735–1744
2019
Later among the works it cites.
R. Hebbar, K. Somandepalli, and S. Narayanan, “Robust speech activity detection in movie audio: Data resources and experimental evaluation,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 4105–4109
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in Proc. IEEE ICASSP 2017 , New Orleans, LA, 2017
2017
Cited alongside, same era.
——, “Objects that sound,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 435–451
2018
Cited alongside, same era.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 631–648
2018
Cited alongside, same era.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 570–586
2018
Cited alongside, same era.
K. Somandepalli, V. Martinez, N. Kumar, and S. Narayanan, “Multimodal representation of advertisements using segment-level autoencoders,” in Proceedings of the 20th ACM International Conference on Multimodal Interaction , 2018, pp. 418–422
2018
Cited alongside, same era.
A. Chattopadhay, A. Sarkar, P. Howlader, and V. N. Balasubramanian, “Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks,” in 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) , 2018, pp. 839–847
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. Sharma, K. Somandepalli, and S. Narayanan, “Toward visual voice activity detection for unconstrained videos,” in 2019 IEEE International Conference on Image Processing (ICIP) . IEEE, 2019, pp. 2991–2995
2019
Cited alongside, same era.
Y. Wang, J. Li, and F. Metze, “A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 31–35
2019
Later among the works it cites.
H. Xu, R. Zeng, Q. Wu, M. Tan, and C. Gan, “Cross-modal relation-aware networks for audio-visual event localization,” in Proceedings of the 28th ACM International Conference on Multimedia , ser. MM ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 3893–3901. [Online]. Available: https://doi.org/10.1145/3394171.3413581
2020
Closest in time.
J. L. Alcazar, F. Caba, L. Mai, F. Perazzi, J.-Y. Lee, P. Arbelaez, and B. Ghanem, “Active speakers in context,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Closest in time.
2020
Closest in time.
W. Wang, D. Tran, and M. Feiszli, “What makes training multi-modal classification networks hard?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 12 695–12 705
2020
Closest in time.
Z. Ren, Z. Yu, X. Yang, M.-Y. Liu, Y. J. Lee, A. G. Schwing, and J. Kautz, “Instance-aware, context-focused, and memory-efficient weakly supervised object detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 10 598–10 607
2020
Closest in time.
2020
Closest in time.
K. Somandepalli, T. Guha, V. R. Martinez, N. Kumar, H. Adam, and S. Narayanan, “Computational media intelligence: Human-centered machine analysis of media,” Proceedings of the IEEE , pp. 1–20, 2021
2021
Closest in time.
C. Song, N. Ning, Y. Zhang, and B. Wu, “A multimodal fake news detection model based on crossmodal attention residual and multichannel convolutional neural networks,” Information Processing and Management , vol. 58, no. 1, p. 102437, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0306457320309304
2021
Closest in time.