Fetching the paper…
Reading the bibliography…
Attention-based models are appealing for multimodal processing because inputs from multiple modalities can be concatenated and fed to a single backbone network - thus requiring very little fusion engineering.
Sensory modalities are not separate modalities: plasticity and interactions
Shinsuke Shimojo and Ladan Shams · 2001
Earlier work this paper cites.
Cross-modal plasticity: where and how?
Daphne Bavelier and Helen J. Neville · 2002
Earlier work this paper cites.
Is neocortex essentially multisensory
Asif A. Ghazanfar and Charles E. Schroeder · 2006
Earlier work this paper cites.
Benefits of multisensory learning
Ladan Shams and Aaron R. Seitz · 2008
Earlier work this paper cites.
Visuo-haptic multisensory object recognition, categorization, and representation
Simon Lacey and K. Sathian · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
ESC: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Task selectivity as a comprehensive principle for brain organization
Amir Amedi, Shir Hofstetter, Shachar Maidenbaum, and Benedetta Heimler · 2017
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelović and Andrew Zisserman · 2017
Earlier work this paper cites.
Task-specific reorganization of the auditory cortex in deaf humans
Lukasz Bola, Maria Zimmermann, Piotr Mostowski, Katarzyna Jednoróg, Artur Marchewka, Paweł Rutkowski, and Marcin Szwed · 2017
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Earlier work this paper cites.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
SpecAugment: A simple data augmentation method for automatic speech recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le · 2019
Cited alongside, same era.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Slow-fast auditory streams for audio recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2021
Later among the works it cites.
Efficient training of audio transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer · 2021
Later among the works it cites.
Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning
Sangho Lee, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas Breuel, Gal Chechik, and Yale Song · 2021
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al · 2021
Later among the works it cites.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Large scale audiovisual learning of sounds with weakly labeled data
Haytham M Fayek and Anurag Kumar · 2020
Cited alongside, same era.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2020
Cited alongside, same era.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Yuki M. Asano, Ruth Fong, João F. Henriques, Geoffrey Zweig, and Andrea Vedaldi · 2020
Cited alongside, same era.
Later among the works it cites.
Broaden your views for self-supervised video learning, 2021
Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altché, Michal Valko, Jean-Bastien Grill, Aäron van den Oord, and Andrew Zisserman · 2021
Later among the works it cites.
Imagenet-21k pretraining for the masses
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
Everything at once–multi-modal fusion transformer for video retrieval
Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, and Hilde Kuehne · 2021
Later among the works it cites.
Eranns: Efficient residual audio neural networks for audio pattern recognition
Sergey Verbitskiy, Vladimir Berikov, and Viacheslav Vyshegorodtsev · 2021
Later among the works it cites.
Co-training transformer with videos and images improves action recognition
Bowen Zhang, Jiahui Yu, Christopher Fifty, Wei Han, Andrew M Dai, Ruoming Pang, and Fei Sha · 2021
Later among the works it cites.
Joao Carreira, Skanda Koppula, Daniel Zoran, Adria Recasens, Catalin Ionescu, Olivier Henaff, Evan Shelhamer, Relja Arandjelovic, Matt Botvinick, Oriol Vinyals, et al · 2022
Later among the works it cites.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Later among the works it cites.
Masked spectrogram prediction for self-supervised audio pre-training
Dading Chong, Helin Wang, Peilin Zhou, and Qingcheng Zeng · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Perceiver IO: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier J. Hénaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, and João Carreira · 2022
Later among the works it cites.
Play it back: Iterative attention for audio recognition
Alexandros Stergiou and Dima Damen · 2022
Later among the works it cites.
Three things everyone should know about vision transformers
Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Jakob Verbeek, and Hervé Jégou · 2022
Later among the works it cites.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.