Fetching the paper…
Reading the bibliography…
In recent years, there has been a growing emphasis on the intersection of audio, vision, and text modalities, driving forward the advancements in multimodal research.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark, 2016
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Vqa: Visual question answering, 2016
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron Weiss, and Kevin Wilson · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering, 2018
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Earlier work this paper cites.
Audio to body dynamics
Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos, 2018
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Earlier work this paper cites.
Learning individual styles of conversational gesture
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp, 2019
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Dancing to music
Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Cited alongside, same era.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Cited alongside, same era.
Rubi: Reducing unimodal biases in visual question answering, 2020
Remi Cadene, Corentin Dancette, Hedi Ben-younes, Matthieu Cord, and Devi Parikh · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations, 2020
Video background music generation with controllable music transformer
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan · 2021
Later among the works it cites.
Audio-visual event localization via recursive fusion by joint co-attention
Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, and Yan Yan · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners, 2021
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
Robust audio-visual instance discrimination
Pedro Morgado, Ishan Misra, and Nuno Vasconcelos · 2021
Later among the works it cites.
How does it sound?
Kun Su, Xiulong Liu, and Eli Shlizerman · 2021
Later among the works it cites.
Cyclic co-learning of sounding object visual grounding and sound separation, 2021
Yapeng Tian, Di Hu, and Chenliang Xu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Improved baselines with momentum contrastive learning, 2020
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Cited alongside, same era.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B. Tenenbaum, and Antonio Torralba · 2020
Cited alongside, same era.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B Tenenbaum, and Antonio Torralba · 2020
Cited alongside, same era.
Bootstrap your own latent: A new approach to self-supervised learning, 2020
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Cited alongside, same era.
Reducing language biases in visual question answering with visually-grounded question encoder, 2020
Gouthaman KV and Anurag Mittal · 2020
Cited alongside, same era.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Nvasconcelos · 2020
Cited alongside, same era.
Later among the works it cites.
Pano-avqa: Grounded audio-visual question answering on 360
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim · 2021
Later among the works it cites.
Learning to answer questions in dynamic audio-visual scenarios, 2022
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution, 2022
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo · 2022
Later among the works it cites.
Cross-modal background suppression for audio-visual event localization
Yan Xia and Zhou Zhao · 2022
Later among the works it cites.
Avqa: A dataset for audio-visual question answering on videos
Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu · 2022
Later among the works it cites.
Contrastive audio-visual masked autoencoder, 2023
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James Glass · 2023
Closest in time.
Vision transformers are parameter-efficient audio-visual learners, 2023
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius · 2023
Closest in time.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Closest in time.
Audio-visual segmentation, 2023
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2023
Closest in time.