Fetching the paper…
Reading the bibliography…
The combination of audio and vision has long been a topic of interest in the multi-modal community.
Dual-modality seq2seq network for audio-visual event localization
Lin, Y.-B.; Li, Y.-J.; and Wang, Y.-C. F. 2019 · 2006
Earlier work this paper cites.
Making a case for 3d convolutions for object segmentation in videos
Mahadevan, S.; Athar, A.; Ošep, A.; Hennen, S.; Leal-Taixé, L.; and Leibe, B. 2020 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J.; Shelhamer, E.; and Darrell, T. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Milletari, F.; Navab, N.; and Ahmadi, S.-A. 2016 · 2016
Earlier work this paper cites.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Wu, Q.; Wang, P.; Shen, C.; Dick, A.; and Van Den Hengel, A. 2016 · 2016
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Objects that sound
Arandjelovic, R.; and Zisserman, A. 2018 · 2018
Cited alongside, same era.
Hybrid task cascade for instance segmentation
Chen, K.; Pang, J.; Wang, J.; Xiong, Y.; Li, X.; Sun, S.; Feng, W.; Liu, Z.; Shi, J.; Ouyang, W.; et al. 2019 · 2019
Cited alongside, same era.
Panoptic feature pyramid networks
Kirillov, A.; Girshick, R.; He, K.; and Dollár, P. 2019 · 2019
Cited alongside, same era.
Upsnet: A unified panoptic segmentation network
Xiong, Y.; Liao, R.; Zhao, H.; Hu, R.; Bai, M.; Yumer, E.; and Urtasun, R. 2019 · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Transformer transforms salient object detection and camouflaged object detection
Mao, Y.; Zhang, J.; Wan, Z.; Dai, Y.; Li, A.; Lv, Y.; Tian, X.; Fan, D.-P.; and Barnes, N. 2021 · 2021
Later among the works it cites.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Wu, Y.; and Yang, Y. 2021 · 2021
Later among the works it cites.
Associating objects with transformers for video object segmentation
Yang, Z.; Wei, Y.; and Yang, Y. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2020
Cited alongside, same era.
Audiovisual transformer with instance attention for audio-visual event localization
Lin, Y.-B.; and Wang, Y.-C. F. 2020 · 2020
Cited alongside, same era.
Multiple sound sources localization from coarse to fine
Qian, R.; Hu, D.; Dinkel, H.; Wu, M.; Xu, N.; and Lin, W. 2020 · 2020
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Tian, Y.; Li, D.; and Xu, C. 2020 · 2020
Cited alongside, same era.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2020
Cited alongside, same era.
Localizing visual sounds the hard way
Chen, H.; Xie, W.; Afouras, T.; Nagrani, A.; Vedaldi, A.; and Zisserman, A. 2021 · 2021
Cited alongside, same era.
Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency Prediction
Zhang, J.; Xie, J.; Barnes, N.; and Li, P. 2021 · 2021
Later among the works it cites.
Vision transformer adapter for dense predictions
Chen, Z.; Duan, Y.; Wang, W.; He, J.; Lu, T.; Dai, J.; and Qiao, Y. 2022 · 2022
Later among the works it cites.
Masked-attention mask transformer for universal image segmentation
Cheng, B.; Misra, I.; Schwing, A. G.; Kirillov, A.; and Girdhar, R. 2022 · 2022
Later among the works it cites.
OneFormer: One Transformer to Rule Universal Image Segmentation
Jain, J.; Li, J.; Chiu, M.; Hassani, A.; Orlov, N.; and Shi, H. 2022 · 2022
Later among the works it cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L. M.; and Shum, H.-Y. 2022 · 2022
Later among the works it cites.
Audio–Visual Segmentation
Zhou, J.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; and Zhong, Y. 2022 · 2022
Later among the works it cites.
Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks
Zhu, X.; Zhu, J.; Li, H.; Wu, X.; Li, H.; Wang, X.; and Dai, J. 2022 · 2022
Later among the works it cites.
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
Mao, Y.; Zhang, J.; Xiang, M.; Lv, Y.; Zhong, Y.; and Dai, Y. 2023 · 2023
Closest in time.
Audio-Visual Segmentation with Semantics
Zhou, J.; Shen, X.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; et al. 2023 · 2023
Closest in time.