Fetching the paper…
Reading the bibliography…
Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment.
Audio-visual sound separation via hidden markov models
John Hershey and Michael Casey · 2001
Earlier work this paper cites.
Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria
Tuomas Virtanen · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Singing-voice separation from monaural recordings using robust principal component analysis
Po-Sen Huang, Scott Deeann Chen, Paris Smaragdis, and Mark Hasegawa-Johnson · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H. McDermott, William T. Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360 video
Pedro Morgado, Nuno Vasconcelos, Timothy Langlois, and Oliver Wang · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A. Efros · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample · 2019
Earlier work this paper cites.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Cited alongside, same era.
Dual-modality seq2seq network for audio-visual event localization
Yan-Bo Lin, Yu-Jhe Li, and Yu-Chiang Frank Wang · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Ernie: Enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu · 2019
Cited alongside, same era.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Yu Wu and Yi Yang · 2021
Later among the works it cites.
BEit: BERT pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He · 2022
Later among the works it cites.
Mix and localize: Localizing sound sources in mixtures
Xixi Hu, Ziyang Chen, and Andrew Owens · 2022
Later among the works it cites.
Masked feature prediction for self-supervised visual pre-training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan L. Yuille, and Christoph Feichtenhofer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Cited alongside, same era.
Recursive visual sound separation using minus-plus net
Xudong Xu, Bo Dai, and Dahua Lin · 2019
Cited alongside, same era.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Cited alongside, same era.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, and Antonio Torralba · 2020
Cited alongside, same era.
Audiovisual transformer with instance attention for audio-visual event localization
Yan-Bo Lin and Yu-Chiang Frank Wang · 2020
Cited alongside, same era.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Vasconcelos · 2020
Cited alongside, same era.
Jiantao Wu and Shentong Mo · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2022
Later among the works it cites.
Audio-visual segmentation
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2022
Later among the works it cites.
Peco: Perceptual codebook for BERT pre-training of vision transformers
Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu · 2023
Closest in time.
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab · 2023
Closest in time.
Contrastive audio-visual masked autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R. Glass · 2023
Closest in time.
Vision transformers are parameter-efficient audio-visual learners
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius · 2023
Closest in time.
Multimodal variational auto-encoder based audio-visual segmentation
Yuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong, and Yuchao Dai · 2023
Closest in time.
A unified audio-visual learning framework for localization, separation, and recognition
Shentong Mo and Pedro Morgado · 2023
Closest in time.
Weakly-supervised audio-visual segmentation
Shentong Mo and Bhiksha Raj · 2023
Closest in time.
Audio-visual class-incremental learning
Weiguo Pian, Shentong Mo, Yunhui Guo, and Yapeng Tian · 2023
Closest in time.
Should you mask 15% in masked language modeling?
Alexander Wettig, Tianyu Gao, Zexuan Zhong, and Danqi Chen · 2023
Closest in time.
Audio-visual segmentation with semantics
Jinxing Zhou, Xuyang Shen, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2023
Closest in time.