Fetching the paper…
Reading the bibliography…
Recent Audio-Visual Question Answering (AVQA) methods rely on complete visual and audio input to answer questions accurately.
Raij, T., Uutela, K., Hari, R.: Audiovisual integration of letters in the human brain. Neuron (2000)
2000
Earlier work this paper cites.
Calvert, G.A., Hansen, P.C., Iversen, S.D., Brammer, M.J.: Detection of audio-visual integration sites in humans by application of electrophysiological criteria to the bold effect. Neuroimage (2001)
2001
Earlier work this paper cites.
Sweller, J.: Instructional design consequences of an analogy between evolution by natural selection and human cognitive architecture. Instructional science (2004)
2004
Earlier work this paper cites.
McGrew, K.S.: Chc theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research (2009)
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
Lindenberger, U.: Human cognitive aging: corriger la fortune? science (2014)
2014
Earlier work this paper cites.
Lahat, D., Adali, T., Jutten, C.: Multimodal data fusion: an overview of methods, challenges, and prospects. Proceedings of the IEEE (2015)
2015
Earlier work this paper cites.
Cai, L., Wang, Z., Gao, H., Shen, D., Ji, S.: Deep adversarial learning for multi-modality missing data completion. In: Int. Conf. Knowledge Discovery and Data Mining (2018)
2018
Earlier work this paper cites.
Schwartz, I., Schwing, A.G., Hazan, T.: A simple baseline for audio-visual scene-aware dialog. In: IEEE Conf. Comput. Vis. Pattern Recog. (2019)
2019
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Adv. Neural Inform. Process. Syst. (2020)
2020
Earlier work this paper cites.
Kim, D., Park, K., Lee, G.: Oddeyecam: A sensing technique for body-centric peephole interaction using wfov rgb and nfov depth cameras. In: ACM Symp. User Interface Software Technology (2020)
2020
Earlier work this paper cites.
Parthasarathy, S., Sundaram, S.: Training strategies to handle missing modalities for audio-visual expression recognition. In: Proc. ACM Int. Conf. Multimodal Interact. (2020)
2020
Earlier work this paper cites.
Chen, Y., Xian, Y., Koepke, A., Shan, Y., Akata, Z.: Distilling audio-visual knowledge by compositional contrastive learning. In: IEEE Conf. Comput. Vis. Pattern Recog. (2021)
2021
Earlier work this paper cites.
Choi, C., Choi, J.H., Li, J., Malla, S.: Shared cross-modal trajectory prediction for autonomous driving. In: IEEE Conf. Comput. Vis. Pattern Recog. (2021)
2021
Earlier work this paper cites.
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. Adv. Neural Inform. Process. Syst. (2021)
2021
Earlier work this paper cites.
Ma, M., Ren, J., Zhao, L., Tulyakov, S., Wu, C., Peng, X.: Smil: Multimodal learning with severely missing modality. In: AAAI (2021)
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. (2021)
2021
Cited alongside, same era.
Yun, H., Yu, Y., Yang, W., Lee, K., Kim, G.: Pano-avqa: Grounded audio-visual question answering on 360deg videos. In: Int. Conf. Comput. Vis. (2021)
2021
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Later among the works it cites.
Saharia, C., Chan, W., Chang, H., Lee, C., Ho, J., Salimans, T., Fleet, D., Norouzi, M.: Palette: Image-to-image diffusion models. In: ACM SIGGRAPH (2022)
2022
Later among the works it cites.
Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D.J., Norouzi, M.: Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. (2022)
2022
Later among the works it cites.
Yang, P., Wang, X., Duan, X., Chen, H., Hou, R., Jin, C., Zhu, W.: Avqa: A dataset for audio-visual question answering on videos. In: ACM Int. Conf. Multimedia (2022)
2022
Later among the works it cites.
Jin, T., Cheng, X., Li, L., Lin, W., Wang, Y., Zhao, Z.: Rethinking missing modality learning from a decoding perspective. In: ACM Int. Conf. Multimedia (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeng, Y., Yang, H., Chao, H., Wang, J., Fu, J.: Improving visual quality of image synthesis by a token-based generator with transformers. Adv. Neural Inform. Process. Syst. (2021)
2021
Cited alongside, same era.
Zhang, J., Xu, X., Shen, F., Lu, H., Liu, X., Shen, H.T.: Enhancing audio-visual association with self-supervised curriculum learning. In: AAAI (2021)
2021
Cited alongside, same era.
Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., Guo, B.: Vector quantized diffusion model for text-to-image synthesis. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Cited alongside, same era.
Kawar, B., Elad, M., Ermon, S., Song, J.: Denoising diffusion restoration models. Adv. Neural Inform. Process. Syst. (2022)
2022
Cited alongside, same era.
Kim, G., Kwon, T., Ye, J.C.: Diffusionclip: Text-guided diffusion models for robust image manipulation. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Cited alongside, same era.
Kim, J.U., Park, S., Ro, Y.M.: Towards versatile pedestrian detector with multisensory-matching and multispectral recalling memory. In: AAAI (2022)
2022
Cited alongside, same era.
Lee, S., Kim, H.I., Ro, Y.M.: Weakly paired associative learning for sound and image representations via bimodal associative memory. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Cited alongside, same era.
Lee, S., Park, S., Ro, Y.M.: Audio-visual mismatch-aware video retrieval via association and adjustment. In: Eur. Conf. Comput. Vis. (2022)
2022
Cited alongside, same era.
2023
Later among the works it cites.
Kim, J.U., Ro, Y.M.: Enabling visual object detection with object sounds via visual modality recalling memory. IEEE Trans. Neural Netw. Learn. Syst. (2023)
2023
Later among the works it cites.
Lee, Y.L., Tsai, Y.H., Chiu, W.C., Lee, C.Y.: Multimodal prompting with missing modalities for visual recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Later among the works it cites.
Li, G., Hou, W., Hu, D.: Progressive spatio-temporal perception for audio-visual question answering. In: ACM Int. Conf. Multimedia (2023)
2023
Later among the works it cites.
Pian, W., Mo, S., Guo, Y., Tian, Y.: Audio-visual class-incremental learning. In: Int. Conf. Comput. Vis. (2023)
2023
Later among the works it cites.
Um, S.J., Kim, D., Kim, J.U.: Audio-visual spatial integration and recursive attention for robust sound source localization. In: ACM Int. Conf. Multimedia (2023)
2023
Later among the works it cites.
Wang, H., Chen, Y., Ma, C., Avery, J., Hull, L., Carneiro, G.: Multi-modal learning with missing modality via shared-specific feature modelling. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Later among the works it cites.
Woo, S., Lee, S., Park, Y., Nugroho, M.A., Kim, C.: Towards good practices for missing modality robust action recognition. In: AAAI (2023)
2023
Later among the works it cites.
Kim, D., Um, S.J., Lee, S., Kim, J.U.: Learning to visually localize sound sources from mixtures without prior source knowledge. In: IEEE Conf. Comput. Vis. Pattern Recog. (2024)
2024
Closest in time.
Maheshwari, H., Liu, Y.C., Kira, Z.: Missing modality robustness in semi-supervised multi-modal semantic segmentation. In: IEEE Winter Conf. Appl. Comput. Vis. (2024)
2024
Closest in time.
Wu, R., Wang, H., Dayoub, F., Chen, H.T.: Segment beyond view: Handling partially missing modality for audio-visual semantic segmentation. In: AAAI (2024)
2024
Closest in time.
Yao, W., Yin, K., Cheung, W.K., Liu, J., Qin, J.: Drfuse: Learning disentangled representation for clinical multi-modal fusion with missing modality and modal inconsistency. In: AAAI (2024)
2024
Closest in time.