Fetching the paper…
Reading the bibliography…
Audio-visual generalized zero-shot learning is a rapidly advancing domain that seeks to understand the intricate relations between audio and visual cues within videos.
Tammes, P.M.L.: On the origin of number and arrangement of the places of exit on the surface of pollen-grains. Recueil des travaux botaniques néerlandais 27
1930
Earlier work this paper cites.
Hershey, J., Casey, M.: Audio-visual sound separation via hidden markov models. Advances in Neural Information Processing Systems 14
2001
Earlier work this paper cites.
2012
Earlier work this paper cites.
Mikolov, T., Chen, K., Corrado, G., Dean, J.: Efficient estimation of word representations in vector space. In: Proceedings of International Conference on Learning Representations (ICLR) (2013)
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
Heilbron, F.C., Escorcia, V., Ghanem, B., Niebles, J.C.: Activitynet: A large-scale video benchmark for human activity understanding. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 961–970 (2015)
2015
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV) (2015)
2015
Earlier work this paper cites.
Aytar, Y., Vondrick, C., Torralba, A.: Soundnet: Learning sound representations from unlabeled video. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS) (2016)
2016
Earlier work this paper cites.
Owens, A., Wu, J., McDermott, J.H., Freeman, W.T., Torralba, A.: Ambient sound provides supervision for visual learning. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 801–816 (2016)
2016
Earlier work this paper cites.
Arandjelovic, R., Zisserman, A.: Look, listen and learn. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). pp. 609–617 (2017)
2017
Earlier work this paper cites.
Hershey, S., Chaudhuri, S., Ellis, D.P.W., Gemmeke, J.F., Jansen, A., Moore, R.C., Plakal, M., Platt, D., Saurous, R.A., Seybold, B., Slaney, M., Weiss, R.J., Wilson, K.: Cnn architectures for large-scale audio classification. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2017)
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Korbar, B., Tran, D., Torresani, L.: Cooperative learning of audio and video models from self-supervised synchronization. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS) (2018)
2018
Earlier work this paper cites.
Morgado, P., Vasconcelos, N., Langlois, T., Wang, O.: Self-supervised generation of spatial audio for 360 video. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS) (2018)
2018
Earlier work this paper cites.
Senocak, A., Oh, T.H., Kim, J., Yang, M.H., Kweon, I.S.: Learning to localize sound source in visual scenes. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4358–4366 (2018)
2018
Earlier work this paper cites.
Tian, Y., Shi, J., Li, B., Duan, Z., Xu, C.: Audio-visual event localization in unconstrained videos. In: Proceedings of European Conference on Computer Vision (ECCV) (2018)
2018
Earlier work this paper cites.
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., Torralba, A.: The sound of pixels. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 570–586 (2018)
2018
Earlier work this paper cites.
Gao, R., Grauman, K.: 2.5d visual sound. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 324–333 (2019)
2019
Earlier work this paper cites.
Hu, D., Nie, F., Li, X.: Deep multimodal clustering for unsupervised audiovisual learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 9248–9257 (2019)
2019
Earlier work this paper cites.
Lin, Y.B., Li, Y.J., Wang, Y.C.F.: Dual-modality seq2seq network for audio-visual event localization. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 2002–2006 (2019)
2019
Earlier work this paper cites.
Wu, Y., Zhu, L., Yan, Y., Yang, Y.: Dual attention matching for audio-visual event localization. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). pp. 6291–6299 (2019)
2019
Earlier work this paper cites.
Zhao, H., Gan, C., Ma, W.C., Torralba, A.: The sound of motions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 1735–1744 (2019)
2019
Cited alongside, same era.
Chen, C., Jain, U., Schissler, C., Garí, S.V.A., Al-Halah, Z., Ithapu, V.K., Robinson, P., Grauman, K.: Soundspaces: Audio-visual navigation in 3d environments. In: Proceedings of European Conference on Computer Vision (ECCV). pp. 17–36 (2020)
2020
Cited alongside, same era.
Chen, H., Xie, W., Vedaldi, A., Zisserman, A.: Vggsound: A large-scale audio-visual dataset. In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 721–725. IEEE (2020)
2020
Cited alongside, same era.
Gan, C., Huang, D., Zhao, H., Tenenbaum, J.B., Torralba, A.: Music gesture for visual sound separation. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10478–10487 (2020)
2020
Cited alongside, same era.
Mercea, O.B., Hummel, T., Koepke, A.S., Akata, Z.: Temporal and cross-modal attention for audio-visual zero-shot learning. In: Proceedings of European Conference on Computer Vision (ECCV) (2022)
2022
Later among the works it cites.
Mercea, O.B., Riesch, L., Koepke, A.S., Akata, Z.: Audio-visual generalised zero-shot learning with cross-modal attention and language. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10553–10563 (2022)
2022
Later among the works it cites.
Mo, S., Morgado, P.: Benchmarking weakly-supervised audio-visual sound localization. In: European Conference on Computer Vision (ECCV) Workshop (2022)
2022
Later among the works it cites.
Mo, S., Morgado, P.: A closer look at weakly-supervised audio-visual source localization. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, Y.B., Wang, Y.C.F.: Audiovisual transformer with instance attention for audio-visual event localization. In: Proceedings of the Asian Conference on Computer Vision (ACCV) (2020)
2020
Cited alongside, same era.
Morgado, P., Li, Y., Vasconcelos, N.: Learning representations from audio-visual spatial alignment. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS). pp. 4733–4744 (2020)
2020
Cited alongside, same era.
Morgado, P., Li, Y., Costa Pereira, J., Saberian, M., Vasconcelos, N.: Deep hashing with hash-consistent large margin proxy embeddings. International Journal of Computer Vision (2020)
2020
Cited alongside, same era.
Parida, K.K., Matiyali, N., Guha, T., Sharma, G.: Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classification and retrieval of videos. In: Proceedings of 2020 IEEE Winter Conference on Applications of Computer Vision (WACV). pp. 3240–3249 (2020)
2020
Cited alongside, same era.
Tian, Y., Li, D., Xu, C.: Unified multisensory perception: Weakly-supervised audio-visual video parsing. In: Proceedings of European Conference on Computer Vision (ECCV). p. 436–454 (2020)
2020
Cited alongside, same era.
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the International Conference on Computer Vision (ICCV) (2021)
2021
Cited alongside, same era.
Chen, C., Majumder, S., Ziad, A.H., Gao, R., Kumar Ramakrishnan, S., Grauman, K.: Learning to set waypoints for audio-visual navigation. In: Proceedings of International Conference on Learning Representations (ICLR) (2021)
2021
Cited alongside, same era.
Fayek, H.M., Kumar, A.: Large scale audiovisual learning of sounds with weakly labeled data. In: Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (2021)
2021
Cited alongside, same era.
Mo, S., Tian, Y.: Multi-modal grouping network for weakly-supervised audio-visual video parsing. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Later among the works it cites.
Mo, S., Tian, Y.: Semantic-aware multi-modal grouping for weakly-supervised audio-visual video parsing. In: European Conference on Computer Vision (ECCV) Workshop (2022)
2022
Later among the works it cites.
Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K.V., Joulin, A., Misra, I.: Imagebind: One embedding space to bind them all. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023)
2023
Later among the works it cites.
Hong, J., Hayder, Z., Han, J., Fang, P., Harandi, M., Petersson, L.: Hyperbolic audio-visual zero-shot learning. In: Proceedings of the International Conference on Computer Vision (ICCV) (2023)
2023
Later among the works it cites.
Mo, S., Morgado, P.: A unified audio-visual learning framework for localization, separation, and recognition. In: Proceedings of the International Conference on Machine Learning (ICML) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., Dubnov, S.: Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In: IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP (2023)
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.