Fetching the paper…
Reading the bibliography…
Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
VGGSound: A Large-scale Audio-Visual Dataset
Chen, H.; Xie, W.; Vedaldi, A.; and Zisserman, A. 2020 · 2004
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Earlier work this paper cites.
Objects that sound
Arandjelovic, R.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Baltrušaitis, T.; Ahuja, C.; and Morency, L.-P. 2018 · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Senocak, A.; Oh, T.-H.; Kim, J.; Yang, M.-H.; and Kweon, I. S. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Localizing visual sounds the hard way
Chen, H.; Xie, W.; Afouras, T.; Nagrani, A.; Vedaldi, A.; and Zisserman, A. 2021 · 2021
Earlier work this paper cites.
Class-aware sounding objects localization via audiovisual correspondence
Hu, D.; Wei, Y.; Qian, R.; Lin, W.; Song, R.; and Wen, J.-R. 2021 · 2021
Cited alongside, same era.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Cited alongside, same era.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference
Schick, T.; and Schütze, H. 2021 · 2021
Cited alongside, same era.
Prompt-based Distribution Alignment for Domain Generalization in Text Classification
Prompt vision transformer for domain generalization
Zheng, Z.; Yue, X.; Wang, K.; and You, Y. 2022 · 2022
Later among the works it cites.
Prompt Consistency for Zero-Shot Task Generalization
Zhou, C.; He, J.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2022a · 2022
Later among the works it cites.
AVSegFormer: Audio-Visual Segmentation with Transformer
Gao, S.; Chen, Z.; Chen, G.; Wang, W.; and Lu, T. 2023 · 2023
Closest in time.
Multimodal Prompt Learning in Emotion Recognition Using Context and Audio Information
Jeong, E.; Kim, G.; and Kang, S. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jia, C.; and Zhang, Y. 2022 · 2022
Cited alongside, same era.
Learning common and specific visual prompts for domain generalization
Li, A.; Zhuang, L.; Fan, S.; and Wang, S. 2022 · 2022
Cited alongside, same era.
Test-time prompt tuning for zero-shot generalization in vision-language models
Shu, M.; Nie, W.; Huang, D.-A.; Yu, Z.; Goldstein, T.; Anandkumar, A.; and Xiao, C. 2022 · 2022
Cited alongside, same era.
Learning in audio-visual context: A review, analysis, and new perspective
Wei, Y.; Hu, D.; Tian, Y.; and Li, X. 2022 · 2022
Cited alongside, same era.
Unified vision and language prompt learning
Zang, Y.; Li, W.; Zhou, K.; Huang, C.; and Loy, C. C. 2022 · 2022
Cited alongside, same era.
Audio-Visual Segmentation by Exploring Cross-Modal Mutual Semantics
Liu, C.; Li, P.; Qi, X.; Zhang, H.; Li, L.; Wang, D.; and Yu, X. 2023a
Cited in the paper.
Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation
Liu, J.; Ju, C.; Ma, C.; Wang, Y.; Wang, Y.; and Zhang, Y. 2023b
Cited in the paper.
Hear to Segment: Unmixing the Audio to Guide the Semantic Segmentation
Ling, Y.; Li, Y.; Gan, Z.; Zhang, J.; Chi, M.; and Wang, Y. 2023 · 2023
Closest in time.
AV-SAM: Segment anything model meets audio-visual localization and segmentation
Mo, S.; and Tian, Y. 2023 · 2023
Closest in time.
Marginnce: Robust sound localization with a negative margin
Park, S.; Senocak, A.; and Chung, J. S. 2023 · 2023
Closest in time.
Foundation models for decision making: Problems, methods, and opportunities
Yang, S.; Nachum, O.; Du, Y.; Wei, J.; Abbeel, P.; and Schuurmans, D. 2023 · 2023
Closest in time.