Fetching the paper…
Reading the bibliography…
Traditional audio-visual methods rely on independent audio and visual backbones, which is costly and not scalable.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A. Efros · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Earlier work this paper cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Earlier work this paper cites.
Labelling unlabelled videos from scratch with multi-modal self-supervision
Yuki M Asano, Mandela Patrick, Christian Rupprecht, and Andrea Vedaldi · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Large scale audiovisual learning of sounds with weakly labeled data
Haytham M Fayek and Anurag Kumar · 2020
Earlier work this paper cites.
Audiovisual transformer with instance attention for audio-visual event localization
Yan-Bo Lin and Yu-Chiang Frank Wang · 2020
Earlier work this paper cites.
What makes the sound?: A dual-modality interacting network for audio-visual event localization
Janani Ramaswamy · 2020
Earlier work this paper cites.
See the sound, hear the pixels
Janani Ramaswamy and Sukhendu Das · 2020
Earlier work this paper cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Earlier work this paper cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Earlier work this paper cites.
Cross-modal attention network for temporal inconsistent audio-visual event localization
Hanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang, and Yan Yan · 2020
Earlier work this paper cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Earlier work this paper cites.
Distilling audio-visual knowledge by compositional contrastive learning
Yanbei Chen, Yongqin Xian, A Koepke, Ying Shan, and Zeynep Akata · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Earlier work this paper cites.
Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation
Yuan Gong, Yu-An Chung, and James Glass · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira · 2021
Cited alongside, same era.
Cross-attentional audio-visual fusion for weakly-supervised action localization
Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, and Sungrack Yun · 2021
Cited alongside, same era.
Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning
Sangho Lee, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas Breuel, Gal Chechik, and Yale Song · 2021
Cited alongside, same era.
Mst: Masked self-supervised transformer for visual representation
Zhaowen Li, Zhiyang Chen, Fan Yang, Wei Li, Yousong Zhu, Chaoyang Zhao, Rui Deng, Liwei Wu, Rui Zhao, Ming Tang, et al · 2021
Cited alongside, same era.
Polyvit: Co-training vision transformers on images, videos and audio
Valerii Likhosherstov, Anurag Arnab, Krzysztof Choromanski, Mario Lucic, Yi Tay, Adrian Weller, and Mostafa Dehghani · 2021
Cited alongside, same era.
Learning audio-visual speech representation by masked multimodal cluster prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed · 2022
Later among the works it cites.
Everything at once–multi-modal fusion transformer for video retrieval
Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, and Hilde Kuehne · 2022
Later among the works it cites.
Tvlt: Textless vision-language transformer
Zineng Tang, Jaemin Cho, Yixin Nie, and Mohit Bansal · 2022
Later among the works it cites.
Evading the simplicity bias: Training a diverse set of models discovers solutions with superior ood generalization
Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton Van den Hengel · 2022
Later among the works it cites.
Equivariance and invariance inductive bias for learning from insufficient data
Tan Wad, Qianru Sun, Sugiri Pranata, Karlekar Jayashree, and Hanwang Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang · 2021
Cited alongside, same era.
Active contrastive learning of audio-visual video representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2021
Cited alongside, same era.
Contrastive learning of global and local audio-visual representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2021
Cited alongside, same era.
Robust audio-visual instance discrimination
Pedro Morgado, Ishan Misra, and Nuno Vasconcelos · 2021
Cited alongside, same era.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2021
Cited alongside, same era.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
Cited alongside, same era.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Yu Wu and Yi Yang · 2021
Cited alongside, same era.
Masked feature prediction for self-supervised visual pre-training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer · 2022
Later among the works it cites.
Cross-modal background suppression for audio-visual event localization
Yan Xia and Zhou Zhao · 2022
Later among the works it cites.
Masked autoencoders that listen
Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, Christoph Feichtenhofer, et al · 2022
Later among the works it cites.
Learning visual representation from modality-shared contrastive language-image pre-training
Haoxuan You, Luowei Zhou, Bin Xiao, Noel Codella, Yu Cheng, Ruochen Xu, Shih-Fu Chang, and Lu Yuan · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Audio-adaptive activity recognition across video domains
Yunhua Zhang, Hazel Doughty, Ling Shao, and Cees GM Snoek · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al · 2023
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab · 2023
Later among the works it cites.
Omnimae: Single model masked pretraining on images and videos
Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Contrastive audio-visual masked autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James Glass · 2023
Later among the works it cites.
Mavil: Masked audio-video learners
Po-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali, Haoqi Fan, Yanghao Li, Shang-Wen Li, Gargi Ghosh, Jitendra Malik, and Christoph Feichtenhofer · 2023
Later among the works it cites.
Audio-visual contrastive learning with temporal self-supervision
Simon Jenni, Alexander Black, and John Collomosse · 2023
Later among the works it cites.
Scaling language-image pre-training via masking
Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He · 2023
Later among the works it cites.
Vision transformers are parameter-efficient audio-visual learners
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius · 2023
Later among the works it cites.
Ave-clip: Audioclip-based multi-window temporal transformer for audio visual event localization
Tanvir Mahmud and Diana Marculescu · 2023
Later among the works it cites.
A unified audio-visual learning framework for localization, separation, and recognition
Shentong Mo and Pedro Morgado · 2023
Later among the works it cites.
Class-incremental grouping network for continual audio-visual learning
Shentong Mo, Weiguo Pian, and Yapeng Tian · 2023
Later among the works it cites.
Zorro: the masked multimodal transformer
Adrià Recasens, Jason Lin, Joāo Carreira, Drew Jaegle, Luyu Wang, Jean-baptiste Alayrac, Pauline Luc, Antoine Miech, Lucas Smaira, Ross Hemsley, et al · 2023
Later among the works it cites.
Clippo: Image-and-language understanding from pixels only
Michael Tschannen, Basil Mustafa, and Neil Houlsby · 2023
Later among the works it cites.
One-peace: Exploring one general representation model toward unlimited modalities
Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai, Xiaohuan Zhou, Jingren Zhou, Xinggang Wang, and Chang Zhou · 2023
Later among the works it cites.
Contrastive learning relies more on spatial inductive bias than supervised learning: An empirical study
Yuanyi Zhong, Haoran Tang, Jun-Kun Chen, and Yu-Xiong Wang · 2023
Later among the works it cites.