Fetching the paper…
Reading the bibliography…
Sight and hearing are two senses that play a vital role in human communication and scene understanding.
B. Jones and B. Kabanoff, “Eye movements in auditory space perception,” Perception & Psychophysics , 1975
1975
Earlier work this paper cites.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature , 1976
1976
Earlier work this paper cites.
D. R. Reddy, “Speech recognition by machine: A review,” Proceedings of the IEEE , 1976
1976
Earlier work this paper cites.
K. Hikosaka, E. Iwai, H. a. Saito, and K. Tanaka, “Polysensory properties of neurons in the anterior bank of the caudal superior temporal sulcus of the macaque monkey,” Journal of neurophysiology , 1988
1988
Earlier work this paper cites.
V. R. de Sa, “Learning classification with unlabeled data,” NeurIPS , 1994
1994
Earlier work this paper cites.
J. Luettin, N. A. Thacker, and S. W. Beet, “Visual speech recognition using active shape models and hidden markov models,” in ICASSP , 1996
1996
Earlier work this paper cites.
L. G. Cohen, P. Celnik, A. Pascual-Leone, B. Corwell, L. Faiz, J. Dambrosia, M. Honda, N. Sadato, C. Gerloff, M. Hallett et al. , “Functional relevance of cross-modal plasticity in blind humans,” Nature , 1997
1997
Earlier work this paper cites.
S. Ben-Yacoub, Y. Abdeljaoued, and E. Mayoraz, “Fusion of face and speech data for person identity verification,” IEEE transactions on neural networks , 1999
1999
Earlier work this paper cites.
T. Choudhury, B. Clarkson, T. Jebara, and A. Pentland, “Multimodal person recognition using unconstrained audio and video,” in AVBPA , 1999
1999
Earlier work this paper cites.
J. Hershey and J. Movellan, “Audio vision: Using audio-visual synchrony to locate sounds,” NeurIPS , 1999
1999
Earlier work this paper cites.
S. Dupont and J. Luettin, “Audio-visual speech modeling for continuous speech recognition,” IEEE transactions on multimedia , 2000
2000
Earlier work this paper cites.
J. W. Fisher III, T. Darrell, W. Freeman, and P. Viola, “Learning joint statistical models for audio-visual fusion and segregation,” NeurIPS , 2000
2000
Earlier work this paper cites.
P. L. Lai and C. Fyfe, “Kernel and nonlinear canonical correlation analysis,” International Journal of Neural Systems , 2000
2000
Earlier work this paper cites.
T. Chen, “Audiovisual speech processing,” IEEE signal processing magazine , 2001
2001
Earlier work this paper cites.
I. Matthews, G. Potamianos, C. Neti, and J. Luettin, “A comparison of model and transform-based visual features for audio-visual lvcsr,” in ICME , 2001
2001
Earlier work this paper cites.
J. Hershey and M. Casey, “Audio-visual sound separation via hidden markov models,” NeurIPS , 2001
2001
Earlier work this paper cites.
L. Girin, J.-L. Schwartz, and G. Feng, “Audio-visual enhancement of speech in noise,” The Journal of the Acoustical Society of America , 2001
2001
Earlier work this paper cites.
A. V. Nefian, L. Liang, X. Pi, X. Liu, and K. Murphy, “Dynamic bayesian networks for audio-visual speech recognition,” EURASIP Journal on Advances in Signal Processing , 2002
2002
Earlier work this paper cites.
D. A. Reynolds, “An overview of automatic speaker recognition technology,” in ICASSP , 2002
2002
Earlier work this paper cites.
B. Fasel and J. Luettin, “Automatic facial expression analysis: a survey,” Pattern recognition , 2003
2003
Earlier work this paper cites.
B. E. Stein, T. R. Stanford, M. T. Wallace, J. W. Vaughan, and W. Jiang, “Crossmodal spatial interactions in subcortical and cortical circuits,” 2004
2004
Earlier work this paper cites.
J. Cassell, H. H. Vilhjálmsson, and T. Bickmore, “Beat: the behavior expression animation toolkit,” in Life-Like Characters , 2004
2004
Earlier work this paper cites.
E. Kidron, Y. Y. Schechner, and M. Elad, “Pixels that sound,” in CVPR , 2005
2005
Earlier work this paper cites.
Z. Wu, L. Cai, and H. Meng, “Multi-level fusion of audio and visual features for speaker identification,” in International Conference on Biometrics , 2006
2006
Earlier work this paper cites.
H. Vajaria, T. Islam, S. Sarkar, R. Sankar, and R. Kasturi, “Audio segmentation and speaker localization in meeting videos,” in ICPR , 2006
2006
Earlier work this paper cites.
N. Sebe, I. Cohen, T. Gevers, and T. S. Huang, “Emotion recognition based on joint visual and audio cues,” in ICPR , 2006
2006
Earlier work this paper cites.
M. N. Schmidt and R. K. Olsson, “Single-channel speech separation using sparse non-negative matrix factorization.” in Interspeech , 2006
2006
Earlier work this paper cites.
B. Merker, “Consciousness without a cerebral cortex: A challenge for neuroscience and medicine,” Behavioral and brain sciences , 2007
2007
Earlier work this paper cites.
M. Pantic and M. S. Bartlett, “Machine analysis of facial expressions,” in Face recognition , 2007
2007
Earlier work this paper cites.
X. Li, D. Tao, S. J. Maybank, and Y. Yuan, “Visual music and musical vision,” Neurocomputing , 2008
2008
Earlier work this paper cites.
M. Gurban, J.-P. Thiran, T. Drugman, and T. Dutoit, “Dynamic modality weighting for multi-stream hmms inaudio-visual speech recognition,” in ICMI , 2008
2008
Earlier work this paper cites.
Z. Zeng, J. Tu, B. M. Pianfetti, and T. S. Huang, “Audio–visual affective expression recognition through multistream fused hmm,” IEEE Transactions on multimedia , 2008
2008
Earlier work this paper cites.
J. Ruesch, M. Lopes, A. Bernardino, J. Hornstein, J. Santos-Victor, and R. Pfeifer, “Multimodal saliency-based bottom-up attention a framework for the humanoid robot icub,” in ICRA , 2008
2008
Earlier work this paper cites.
R. Jafri and H. R. Arabnia, “A survey of face recognition techniques,” journal of information processing systems , 2009
2009
Earlier work this paper cites.
M. E. Sargin, H. Aradhye, P. J. Moreno, and M. Zhao, “Audiovisual celebrity recognition in unconstrained web videos,” in ICASSP , 2009
2009
Earlier work this paper cites.
G. Friedland, H. Hung, and C. Yeo, “Multi-modal speaker diarization of real-world meetings using compressed-domain video features,” in ICASSP , 2009
2009
Earlier work this paper cites.
G. Friedland, C. Yeo, and H. Hung, “Visual speaker localization aided by acoustic models,” in ACM MM , 2009
2009
Earlier work this paper cites.
M. S. Gazzaniga, The cognitive neuroscience of mind: a tribute to Michael S. Gazzaniga , 2010
2010
Earlier work this paper cites.
G. Zamora-López, C. Zhou, and J. Kurths, “Cortical hubs form a module for multisensory integration on top of the hierarchy of cortical networks,” Frontiers in neuroinformatics , 2010
2010
Earlier work this paper cites.
R. Poppe, “A survey on vision-based human action recognition,” Image and vision computing , 2010
2010
Earlier work this paper cites.
A. Metallinou, S. Lee, and S. Narayanan, “Decision level combination of multiple modalities for recognition and analysis of emotional expression,” in ICASSP , 2010
2010
Earlier work this paper cites.
I. Almajai and B. Milner, “Visually derived wiener filters for speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , 2010
2010
Earlier work this paper cites.
A. Noulas, G. Englebienne, and B. J. Krose, “Multimodal speaker diarization,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2011
2011
Earlier work this paper cites.
J.-C. Lin, C.-H. Wu, and W.-L. Wei, “Error weighted semi-coupled hidden markov model for audio-visual emotion recognition,” IEEE Transactions on Multimedia , 2011
2011
Earlier work this paper cites.
G. A. Ramirez, T. Baltrušaitis, and L.-P. Morency, “Modeling latent discriminative dynamic of multi-dimensional affective signals,” in ACII , 2011
2011
Earlier work this paper cites.
D. Jiang, Y. Cui, X. Zhang, P. Fan, I. Ganzalez, and H. Sahli, “Audio visual emotion recognition based on triple-stream dynamic bayesian network models,” in ACII , 2011
2011
Earlier work this paper cites.
R. Fan, S. Xu, and W. Geng, “Example-based automatic music-driven conventional dance motion synthesis,” IEEE transactions on visualization and computer graphics , 2011
2011
Earlier work this paper cites.
F. Ofli, E. Erzin, Y. Yemez, and A. M. Tekalp, “Learn2dance: Learning statistical music-to-dance mappings for choreography synthesis,” IEEE Transactions on Multimedia , 2011
2011
Earlier work this paper cites.
B. Schauerte, B. Kühn, K. Kroschel, and R. Stiefelhagen, “Multimodal saliency-based attention for object-based scene analysis,” in IROS , 2011
2011
Earlier work this paper cites.
A. Metallinou, M. Wollmer, A. Katsamanis, F. Eyben, B. Schuller, and S. Narayanan, “Context-sensitive learning for enhanced audiovisual emotion classification,” IEEE Transactions on Affective Computing , 2012
2012
Earlier work this paper cites.
F. Eyben, S. Petridis, B. Schuller, and M. Pantic, “Audiovisual vocal outburst classification in noisy acoustic conditions,” in ICASSP , 2012
2012
Earlier work this paper cites.
C.-M. Huang and B. Mutlu, “Robot behavior toolkit: generating effective social behaviors for robots,” in HRI , 2012
2012
Earlier work this paper cites.
H. Izadinia, I. Saleemi, and M. Shah, “Multimodal analysis for identification and segmentation of moving-sounding objects,” IEEE Transactions on Multimedia , 2012
2012
Earlier work this paper cites.
C.-H. Wu, J.-C. Lin, and W.-L. Wei, “Two-level hierarchical alignment for semi-coupled hmm-based audiovisual emotion recognition with temporal course,” IEEE Transactions on Multimedia , 2013
2013
Earlier work this paper cites.
O. Rudovic, S. Petridis, and M. Pantic, “Bimodal log-linear regression for fusion of audio and visual features,” in ACM MM , 2013
2013
Earlier work this paper cites.
M. Lee, K. Lee, and J. Park, “Music similarity-based approach to generating dance motion sequence,” Multimedia tools and applications , 2013
2013
Earlier work this paper cites.
G. Andrew, R. Arora, J. Bilmes, and K. Livescu, “Deep canonical correlation analysis,” in ICML , 2013
2013
Earlier work this paper cites.
X. Huang, J. Baker, and R. Reddy, “A historical perspective of speech recognition,” Communications of the ACM , 2014
2014
Earlier work this paper cites.
N. Rasiwasia, D. Mahajan, V. Mahadevan, and G. Aggarwal, “Cluster canonical correlation analysis,” in Artificial intelligence and statistics , 2014
2014
Earlier work this paper cites.
M. Gygli, H. Grabner, H. Riemenschneider, and L. V. Gool, “Creating summaries from user videos,” in ECCV , 2014
2014
Earlier work this paper cites.
A. Pannese, D. Grandjean, and S. Frühholz, “Subcortical processing in auditory communication,” Hearing research , 2015
2015
Earlier work this paper cites.
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata, “Audio-visual speech recognition using deep learning,” Applied Intelligence , 2015
2015
Earlier work this paper cites.
B. Milner and T. Le Cornu, “Reconstructing intelligible audio speech from visual speech features,” Interspeech , 2015
2015
Earlier work this paper cites.
M. Akbari and H. Cheng, “Real-time piano music transcription based on computer vision,” IEEE Transactions on Multimedia , 2015
2015
Earlier work this paper cites.
D. Hu, X. Li et al. , “Temporal multimodal learning in audiovisual speech recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
N. Sarafianos, T. Giannakopoulos, and S. Petridis, “Audio-visual speaker diarization using fisher linear semi-discriminant analysis,” Multimedia Tools and Applications , 2016
2016
Earlier work this paper cites.
C. Wang, H. Yang, and C. Meinel, “Exploring multimodal video representation for action recognition,” in IJCNN , 2016
2016
Earlier work this paper cites.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman, “Visually indicated sounds,” in CVPR , 2016
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” NeurIPS , 2016
2016
Earlier work this paper cites.
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in ECCV , 2016
2016
Earlier work this paper cites.
X. Min, G. Zhai, K. Gu, and X. Yang, “Fixation prediction through multimodal analysis,” ACM Transactions on Multimedia Computing, Communications, and Applications , 2016
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in ACCV , 2016
2016
Earlier work this paper cites.
A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency, “Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages,” IEEE Intelligent Systems , 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
I. D. Gebru, S. Ba, X. Li, and R. Horaud, “Audio-visual speaker diarization based on spatiotemporal bayesian fusion,” IEEE transactions on pattern analysis and machine intelligence , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , 2017
2017
Earlier work this paper cites.
A. Mehrabian, “Communication without words,” in Communication theory , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Le Cornu and B. Milner, “Generating intelligible audio speech from visual speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2017
2017
Earlier work this paper cites.
A. Ephrat and S. Peleg, “Vid2speech: speech reconstruction from silent video,” in ICASSP , 2017
2017
Earlier work this paper cites.
A. Ephrat, T. Halperin, and S. Peleg, “Improved speech reconstruction from silent video,” in ICCV Workshops , 2017
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM Transactions on Graphics , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. Li, D. Hu, and X. Lu, “Image2song: Song retrieval via bridging image content and lyric words,” in ICCV , 2017
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in ICCV , 2017
2017
Earlier work this paper cites.
C. Chen, S. Li, Y. Wang, H. Qin, and A. Hao, “Video saliency detection via spatial-temporal fusion and low-rank coherency diffusion,” IEEE transactions on image processing , 2017
2017
Earlier work this paper cites.
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik, “Cognitive mapping and planning for visual navigation,” in CVPR , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” Interspeech , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017
2017
Earlier work this paper cites.
D. Man and R. Olchawa, “Brain biophysics: perception, consciousness, creativity. brain computer interface (bci),” in International Scientific Conference BCI 2018 Opole , 2018
2018
Earlier work this paper cites.
Y. Li, M. Yang, and Z. Zhang, “A survey of multi-view representation learning,” IEEE transactions on knowledge and data engineering , 2018
2018
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence , 2018
2018
Earlier work this paper cites.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” IEEE transactions on pattern analysis and machine intelligence , 2018
2018
Cited alongside, same era.
G. Sell, K. Duh, D. Snyder, D. Etter, and D. Garcia-Romero, “Audio-visual person recognition in multimedia data from the iarpa janus program,” in ICASSP , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L.-P. Morency, “Memory fusion network for multi-view sequential learning,” in AAAI , 2018
2018
Cited alongside, same era.
S. Alexanderson, G. E. Henter, T. Kucherenko, and J. Beskow, “Style-controllable speech-driven gesture synthesis using normalising flows,” in Computer Graphics Forum , 2020
2020
Later among the works it cites.
T. Kucherenko, P. Jonell, S. van Waveren, G. E. Henter, S. Alexandersson, I. Leite, and H. Kjellström, “Gesticulator: A framework for semantically-aware speech-driven gesture generation,” in ICMI , 2020
2020
Later among the works it cites.
C. Ahuja, D. W. Lee, Y. I. Nakano, and L.-P. Morency, “Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach,” in ECCV , 2020
2020
Later among the works it cites.
K. Su, X. Liu, and E. Shlizerman, “Audeo: Audio generation for a silent performance video,” NeurIPS , 2020
2020
Later among the works it cites.
J. H. Christensen, S. Hornauer, and X. Y. Stella, “Batvision: Learning to see 3d spatial layout with two ears,” in ICRA , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hazarika, S. Poria, A. Zadeh, E. Cambria, L.-P. Morency, and R. Zimmermann, “Conversational memory network for emotion recognition in dyadic dialogue videos,” in ACL , 2018
2018
Cited alongside, same era.
D. Ghosal, M. S. Akhtar, D. Chauhan, S. Poria, A. Ekbal, and P. Bhattacharyya, “Contextual inter-modal attention for multi-modal sentiment analysis,” in EMNLP , 2018
2018
Cited alongside, same era.
J.-C. Hou, S.-S. Wang, Y.-H. Lai, Y. Tsao, H.-W. Chang, and H.-M. Wang, “Audio-visual speech enhancement using multimodal deep convolutional neural networks,” IEEE Transactions on Emerging Topics in Computational Intelligence , 2018
2018
Cited alongside, same era.
T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” Interspeech , 2018
2018
Cited alongside, same era.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,” in ECCV , 2018
2018
Cited alongside, same era.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in ECCV , 2018
2018
Cited alongside, same era.
E. Vincent, T. Virtanen, and S. Gannot, Audio source separation and speech enhancement , 2018
2018
Cited alongside, same era.
A. Gabbay, A. Shamir, and S. Peleg, “Visual speech enhancement,” Interspeech , 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
R. Gao, C. Chen, Z. Al-Halah, C. Schissler, and K. Grauman, “Visualechoes: Spatial image representation learning through echolocation,” in ECCV , 2020
2020
Later among the works it cites.
D. Zeng, Y. Yu, and K. Oyama, “Deep triplet neural networks with cluster-cca for audio-visual cross-modal retrieval,” ACM Transactions on Multimedia Computing, Communications, and Applications , 2020
2020
Later among the works it cites.
Y. Chen, X. Lu, and S. Wang, “Deep cross-modal image–voice retrieval in remote sensing,” IEEE Transactions on Geoscience and Remote Sensing , 2020
2020
Later among the works it cites.
Y. Asano, M. Patrick, C. Rupprecht, and A. Vedaldi, “Labelling unlabelled videos from scratch with multi-modal self-supervision,” NeurIPS , 2020
2020
Later among the works it cites.
P. Morgado, Y. Li, and N. Nvasconcelos, “Learning representations from audio-visual spatial alignment,” NeurIPS , 2020
2020
Later among the works it cites.
H. Alwassel, D. Mahajan, B. Korbar, L. Torresani, B. Ghanem, and D. Tran, “Self-supervised learning by cross-modal audio-video clustering,” NeurIPS , 2020
2020
Later among the works it cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , 2020
2020
Later among the works it cites.
A. Tsiami, P. Koutras, and P. Maragos, “Stavis: Spatio-temporal audiovisual saliency network,” in CVPR , 2020
2020
Later among the works it cites.
X. Min, G. Zhai, J. Zhou, X.-P. Zhang, X. Yang, and X. Guan, “A multimodal saliency model for videos with high audio-visual correspondence,” IEEE Transactions on Image Processing , 2020
2020
Later among the works it cites.
C. Chen, U. Jain, C. Schissler, S. V. A. Gari, Z. Al-Halah, V. K. Ithapu, P. Robinson, and K. Grauman, “Soundspaces: Audio-visual navigation in 3d environments,” in ECCV , 2020
2020
Later among the works it cites.
C. Gan, Y. Zhang, J. Wu, B. Gong, and J. B. Tenenbaum, “Look, listen, and act: Towards audio-visual embodied navigation,” in ICRA , 2020
2020
Later among the works it cites.
C. Chen, S. Majumder, Z. Al-Halah, R. Gao, S. K. Ramakrishnan, and K. Grauman, “Learning to set waypoints for audio-visual navigation,” in ICLR , 2020
2020
Later among the works it cites.
Y. Tian, D. Li, and C. Xu, “Unified multisensory perception: Weakly-supervised audio-visual video parsing,” in ECCV , 2020
2020
Later among the works it cites.
Y.-B. Lin and Y.-C. F. Wang, “Audiovisual transformer with instance attention for audio-visual event localization,” in ACCV , 2020
2020
Later among the works it cites.
J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xi et al. , “Ava active speaker: An audio-visual dataset for active speaker detection,” in ICASSP , 2020
2020
Later among the works it cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large-scale audio-visual dataset,” in ICASSP , 2020
2020
Later among the works it cites.
H. Zhu, M.-D. Luo, R. Wang, A.-H. Zheng, and R. He, “Deep audio-visual learning: A survey,” International Journal of Automation and Computing , 2021
2021
Later among the works it cites.
L. Sarı, K. Singh, J. Zhou, L. Torresani, N. Singhal, and Y. Saraf, “A multi-view approach to audio-visual speaker verification,” in ICASSP , 2021
2021
Later among the works it cites.
Y. Qian, Z. Chen, and S. Wang, “Audio-visual deep neural network for robust person verification,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Later among the works it cites.
R. Panda, C.-F. R. Chen, Q. Fan, X. Sun, K. Saenko, A. Oliva, and R. Feris, “Adamml: Adaptive multi-modal learning for efficient video recognition,” in ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
F. Lv, X. Chen, Y. Huang, L. Duan, and G. Lin, “Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences,” in CVPR , 2021
2021
Later among the works it cites.
M. Bedi, S. Kumar, M. S. Akhtar, and T. Chakraborty, “Multi-modal sarcasm detection and humor classification in code-mixed conversations,” IEEE Transactions on Affective Computing , 2021
2021
Later among the works it cites.
J. Lee, S.-W. Chung, S. Kim, H.-G. Kang, and K. Sohn, “Looking into your speech: Learning cross-modal affinity for audio-visual speech separation,” in CVPR , 2021
2021
Later among the works it cites.
M. Chatterjee, J. Le Roux, N. Ahuja, and A. Cherian, “Visual scene graphs for audio source separation,” in ICCV , 2021
2021
Later among the works it cites.
Y. Tian, D. Hu, and C. Xu, “Cyclic co-learning of sounding object visual grounding and sound separation,” in CVPR , 2021
2021
Later among the works it cites.
C. Kong, B. Chen, W. Yang, H. Li, P. Chen, and S. Wang, “Appearance matters, so does audio: Revealing the hidden face via cross-modality transfer,” IEEE Transactions on Circuits and Systems for Video Technology , 2021
2021
Later among the works it cites.
V. K. Kurmi, V. Bajaj, B. N. Patro, K. Venkatesh, V. P. Namboodiri, and P. Jyothi, “Collaborative learning to generate audio-video jointly,” in ICASSP , 2021
2021
Later among the works it cites.
S. Di, Z. Jiang, S. Liu, Z. Wang, L. Zhu, Z. He, H. Liu, and S. Yan, “Video background music generation with controllable music transformer,” in ACM MM , 2021
2021
Later among the works it cites.
V. Iashin and E. Rahtu, “Taming visually guided sound generation,” BMVC , 2021
2021
Later among the works it cites.
X. Xu, H. Zhou, Z. Liu, B. Dai, X. Wang, and D. Lin, “Visually informed binaural audio generation without binaural audios,” in CVPR , 2021
2021
Later among the works it cites.
Y.-B. Lin and Y.-C. F. Wang, “Exploiting audio-visual consistency with partial supervision for spatial audio generation,” in AAAI , 2021
2021
Later among the works it cites.
L. Li, S. Wang, Z. Zhang, Y. Ding, Y. Zheng, X. Yu, and C. Fan, “Write-a-speaker: Text-based emotional and rhythmic talking-head generation,” in AAAI , 2021
2021
Later among the works it cites.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in CVPR , 2021
2021
Later among the works it cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in CVPR , 2021
2021
Later among the works it cites.
R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “Ai choreographer: Music conditioned 3d dance generation with aist++,” in ICCV , 2021, pp. 13 401–13 412
2021
Later among the works it cites.
K. K. Parida, S. Srivastava, and G. Sharma, “Beyond image to depth: Improving depth prediction using echoes,” in CVPR , 2021
2021
Later among the works it cites.
Y. Yin, H. Shrivastava, Y. Zhang, Z. Liu, R. R. Shah, and R. Zimmermann, “Enhanced audio tagging via multi-to single-modal teacher-student mutual learning,” in AAAI , 2021
2021
Later among the works it cites.
Z. Xue, S. Ren, Z. Gao, and H. Zhao, “Multimodal knowledge expansion,” in ICCV , 2021
2021
Later among the works it cites.
Y. Chen, Y. Xian, A. Koepke, Y. Shan, and Z. Akata, “Distilling audio-visual knowledge by compositional contrastive learning,” in CVPR , 2021
2021
Later among the works it cites.
R. Huang, H. Hu, W. Wu, K. Sawada, M. Zhang, and D. Jiang, “Dance revolution: Long-term dance generation with music via curriculum learning,” ICLR , 2021
2021
Later among the works it cites.
F. R. Valverde, J. V. Hurtado, and A. Valada, “There is more than meets the eye: Self-supervised multi-object detection and tracking with sound by distilling multimodal knowledge,” in CVPR , 2021
2021
Later among the works it cites.
L. Zhang, Z. Chen, and Y. Qian, “Knowledge distillation from multi-modality to single-modality for person verification,” Interspeech , 2021
2021
Later among the works it cites.
P. Morgado, I. Misra, and N. Vasconcelos, “Robust audio-visual instance discrimination,” in CVPR , 2021
2021
Later among the works it cites.
P. Morgado, N. Vasconcelos, and I. Misra, “Audio-visual instance discrimination with cross-modal agreement,” in CVPR , 2021
2021
Later among the works it cites.
B. Chen, A. Rouditchenko, K. Duarte, H. Kuehne, S. Thomas, A. Boggust, R. Panda, B. Kingsbury, R. Feris, D. Harwath, J. Glass, M. Picheny, and S.-F. Chang, “Multimodal clustering networks for self-supervised learning from unlabeled videos,” in ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,” NeurIPS , vol. 34, 2021
2021
Later among the works it cites.
D. Hu, Y. Wei, R. Qian, W. Lin, R. Song, and J. Wen, “Class-aware sounding objects localization via audiovisual correspondence.” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Later among the works it cites.
S. Jain, P. Yarlagadda, S. Jyoti, S. Karthik, R. Subramanian, and V. Gandhi, “Vinet: Pushing the limits of visual modality for audio-visual saliency prediction,” in IROS , 2021
2021
Later among the works it cites.
G. Wang, C. Chen, D.-P. Fan, A. Hao, and H. Qin, “From semantic categories to fixations: A novel weakly-supervised visual-auditory saliency detection approach,” in CVPR , 2021
2021
Later among the works it cites.
C. Chen, Z. Al-Halah, and K. Grauman, “Semantic audio-visual navigation,” in CVPR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Majumder, Z. Al-Halah, and K. Grauman, “Move2hear: Active audio-visual source separation,” in ICCV , 2021
2021
Later among the works it cites.
B. Duan, H. Tang, W. Wang, Z. Zong, G. Yang, and Y. Yan, “Audio-visual event localization via recursive fusion by joint co-attention,” in WACV , 2021
2021
Later among the works it cites.
J. Zhou, L. Zheng, Y. Zhong, S. Hao, and M. Wang, “Positive sample propagation along the audio-visual event line,” in CVPR , 2021
2021
Later among the works it cites.
Y.-B. Lin, H.-Y. Tseng, H.-Y. Lee, Y.-Y. Lin, and M.-H. Yang, “Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing,” NeurIPS , 2021
2021
Later among the works it cites.
Y. Wu and Y. Yang, “Exploring heterogeneous clues for weakly-supervised audio-visual video parsing,” in CVPR , 2021
2021
Later among the works it cites.
S. Geng, P. Gao, M. Chatterjee, C. Hori, J. Le Roux, Y. Zhang, H. Li, and A. Cherian, “Dynamic graph representation learning for video dialog via multi-modal shuffled transformers,” in AAAI , 2021
2021
Later among the works it cites.
H. Yun, Y. Yu, W. Yang, K. Lee, and G. Kim, “Pano-avqa: Grounded audio-visual question answering on 360deg videos,” in ICCV , 2021
2021
Later among the works it cites.
2022
Closest in time.
W. Chai and G. Wang, “Deep vision multimodal learning: Methodology, benchmark, and trend,” Applied Sciences , 2022
2022
Closest in time.
Q. Song, B. Sun, and S. Li, “Multimodal sparse transformer network for audio-visual speech recognition,” IEEE Transactions on Neural Networks and Learning Systems , 2022
2022
Closest in time.
2022
Closest in time.
Z. Sun, Q. Ke, H. Rahmani, M. Bennamoun, G. Wang, and J. Liu, “Human action recognition from various data modalities: A review,” IEEE transactions on pattern analysis and machine intelligence , 2022
2022
Closest in time.
J. Chen and C. M. Ho, “Mm-vit: Multi-modal video transformer for compressed video action recognition,” in WACV , 2022
2022
Closest in time.
S. Alfasly, J. Lu, C. Xu, and Y. Zou, “Learnable irrelevant modality dropout for multimodal action recognition on modality-specific annotated videos,” in CVPR , 2022
2022
Closest in time.
——, “Domain generalization through audio-visual relative norm alignment in first person action recognition,” in WACV , 2022
2022
Closest in time.
Y. Zhang, H. Doughty, L. Shao, and C. G. Snoek, “Audio-adaptive activity recognition across video domains,” in CVPR , 2022
2022
Closest in time.
Z. Kang, M. Sadeghi, R. Horaud, X. Alameda-Pineda, J. Donley, and A. Kumar, “The impact of removing head movements on audio-visual speech enhancement,” in ICASSP , 2022
2022
Closest in time.
X. Peng, Y. Wei, A. Deng, D. Wang, and D. Hu, “Balanced multimodal learning via on-the-fly gradient modulation,” in CVPR , 2022
2022
Closest in time.
R. Mira, K. Vougioukas, P. Ma, S. Petridis, B. W. Schuller, and M. Pantic, “End-to-end video-to-speech synthesis using generative adversarial networks,” IEEE Transactions on Cybernetics , 2022
2022
Closest in time.
K. K. Parida, S. Srivastava, and G. Sharma, “Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal attention,” in WACV , 2022
2022
Closest in time.
Y. Liang, Q. Feng, L. Zhu, L. Hu, P. Pan, and Y. Yang, “Seeg: Semantic energized co-speech gesture generation,” in CVPR , 2022
2022
Closest in time.
G. Irie, T. Shibata, and A. Kimura, “Co-attention-guided bilinear model for echo-based depth estimation,” in ICASSP , 2022
2022
Closest in time.
Y. Huang, J. Zhang, S. Liu, Q. Bao, D. Zeng, Z. Chen, and W. Liu, “Genre-conditioned long-term 3d dance generation driven by music,” in ICASSP , 2022
2022
Closest in time.
S. H. Lee, W. Roh, W. Byeon, S. H. Yoon, C. Kim, J. Kim, and S. Kim, “Sound-guided semantic image manipulation,” in CVPR , 2022
2022
Closest in time.
2022
Closest in time.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “Audioclip: Extending clip to image, text and audio,” in ICASSP , 2022
2022
Closest in time.
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi, “Merlot reserve: Neural script knowledge through vision and language and sound,” in CVPR , 2022
2022
Closest in time.
X. Hu, Z. Chen, and A. Owens, “Mix and localize: Localizing sound sources in mixtures,” in CVPR , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu, “Learning to answer questions in dynamic audio-visual scenarios,” in CVPR , 2022
2022
Closest in time.
A. Shah, S. Geng, P. Gao, A. Cherian, T. Hori, T. K. Marks, J. Le Roux, and C. Hori, “Audio-visual scene-aware dialog and reasoning using audio-visual transformers with joint student-teacher learning,” in ICASSP , 2022
2022
Closest in time.
Y. Xia and Z. Zhao, “Cross-modal background suppression for audio-visual event localization,” in CVPR , 2022
2022
Closest in time.
P. Wang, J. Li, M. Ma, and X. Fan, “Distributed audio-visual parsing based on multimodal transformer and deep joint source channel coding,” in ICASSP , 2022
2022
Closest in time.