Fetching the paper…
Reading the bibliography…
Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video.
Automatic music soundtrack generation for outdoor videos from contextual sensor information
Yi Yu, Zhijie Shen, and Roger Zimmermann · 2012
Earlier work this paper cites.
Advisor: Personalized video soundtrack recommendation by late fusion with heuristic rankings
Rajiv Ratn Shah, Yi Yu, and Roger Zimmermann · 2014
Earlier work this paper cites.
Correlation autoencoder hashing for supervised cross-modal search
Yue Cao, Mingsheng Long, et al · 2016
Earlier work this paper cites.
Multimodal Machine Learning: A Survey and Taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2017
Earlier work this paper cites.
A Survey of Multi-View Representation Learning
Yingming Li, Ming Yang, and Zhongfei Zhang · 2017
Earlier work this paper cites.
Face-voice matching using cross-modal embeddings
Shota Horiguchi, Naoyuki Kanda, and Kenji Nagamatsu · 2018
Earlier work this paper cites.
Associative multichannel autoencoder for multimodal word representation
Shaonan Wang, Jiajun Zhang, and Chengqing Zong · 2018
Earlier work this paper cites.
Audio-visual embedding for cross-modal music video retrieval through supervised deep cca
Donghuo Zeng, Yi Yu, and Keizo Oyama · 2018
Earlier work this paper cites.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Earlier work this paper cites.
A new benchmark and approach for fine-grained cross-media retrieval
Xiangteng He, Yuxin Peng, and Liu Xie · 2019
Earlier work this paper cites.
Large-scale representation learning from visually grounded untranscribed speech
Gabriel Ilharco, Yuan Zhang, and Jason Baldridge · 2019
Earlier work this paper cites.
Language learning using speech to image retrieval
Danny Merkx, Stefan L Frank, and Mirjam Ernestus · 2019
Earlier work this paper cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, et al · 2019
Earlier work this paper cites.
Deep fusion: An attention guided factorized bilinear pooling for audio-video emotion recognition
Yuanyuan Zhang, Zi-Rui Wang, and Jun Du · 2019
Earlier work this paper cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Earlier work this paper cites.
Deep supervised cross-modal retrieval
Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng · 2019
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, et al · 2020
Earlier work this paper cites.
Audio–visual domain adaptation using conditional semi-supervised generative adversarial networks
Christos Athanasiadis, Enrique Hortal, and Stylianos Asteriadis · 2020
Earlier work this paper cites.
Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, and Yuejie Zhang · 2020
Cited alongside, same era.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, et al · 2020
Cited alongside, same era.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, et al · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Curriculum audiovisual learning
Di Hu, Zheng Wang, Haoyi Xiong, Dong Wang, Feiping Nie, and Dejing Dou · 2020
Cited alongside, same era.
Audiovisual transformer with instance attention for audio-visual event localization
Localizing visual sounds the hard way
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2021
Later among the works it cites.
Audio-visual event localization via recursive fusion by joint co-attention
Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, and Yan Yan · 2021
Later among the works it cites.
Sound-to-imagination: Unsupervised crossmodal translation using deep dense network architecture
Leonardo A Fanzeres and Climent Nadeu · 2021
Later among the works it cites.
Exploiting audio-visual consistency with partial supervision for spatial audio generation
Yan-Bo Lin and Yu-Chiang Frank Wang · 2021
Later among the works it cites.
Cross-modal attention consistency for video-audio unsupervised learning
Shaobo Min, Qi Dai, Hongtao Xie, Chuang Gan, Yongdong Zhang, and Jingdong Wang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yan-Bo Lin and Yu-Chiang Frank Wang · 2020
Cited alongside, same era.
Active contrastive learning of audio-visual video representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2020
Cited alongside, same era.
Contrastive self-supervised learning of global-local audio-visual representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2020
Cited alongside, same era.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Vasconcelos · 2020
Cited alongside, same era.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Yuki M Asano, Polina Kuznetsova, et al · 2020
Cited alongside, same era.
See the sound, hear the pixels
Janani Ramaswamy and Sukhendu Das · 2020
Cited alongside, same era.
Avlnet: Learning audio-visual language representations from instructional videos
Andrew Rouditchenko, Angie Boggust, David Harwath, et al · 2020
Cited alongside, same era.
Later among the works it cites.
End-to-end video-to-speech synthesis using generative adversarial networks
Rodrigo Mira, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis, Björn W Schuller, and Maja Pantic · 2021
Later among the works it cites.
Spoken moments: Learning joint audio-visual representations from video descriptions
Mathew Monfort, SouYoung Jin, Alexander Liu, et al · 2021
Later among the works it cites.
Cross-modal music-video recommendation: A study of design choices
Laure Pretet, Gael Richard, and Geoffroy Peeters · 2021
Later among the works it cites.
Robust latent representations via cross-modal translation and alignment
Vandana Rajan, Alessio Brutti, and Andrea Cavallaro · 2021
Later among the works it cites.
Broaden your views for self-supervised video learning
Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, et al · 2021
Later among the works it cites.
Learning explicit and implicit latent common spaces for audio-visual cross-modal retrieval
Donghuo Zeng, Jianming Wu, Gen Hattori, Yi Yu, and Rong Xu · 2021
Later among the works it cites.
Variational autoencoder with cca for audio-visual cross-modal retrieval
Jiwei Zhang, Yi Yu, Suhua Tang, Jianming Wu, and Wei Li · 2021
Later among the works it cites.
Adversarial-metric learning for audio-visual cross-modal matching
Aihua Zheng, Menglan Hu, Bo Jiang, Yan Huang, Yan Yan, and Bin Luo · 2021
Later among the works it cites.
Deep co-attention network for multi-view subspace learning
Lecheng Zheng, Yu Cheng, Hongxia Yang, Nan Cao, and Jingrui He · 2021
Later among the works it cites.
Leveraging category information for single-frame visual sound source separation
Lingyu Zhu and Esa Rahtu · 2021
Later among the works it cites.
Deep Audio-visual Learning: A Survey
Hao Zhu, Hao Zhu, Mandi Luo, Rui Wang, Rui Wang, Aihua Zheng, and Ran He · 2021
Later among the works it cites.
Learning audio-visual correlations from variational cross-modal generation
Ye Zhu, Yu Wu, Hugo Latapie, Yi Yang, and Yan Yan · 2021
Later among the works it cites.