Fetching the paper…
Reading the bibliography…
In this work, we propose an efficient Video-Language Alignment (ViLA) network.
MacFarland, S.: If a picture is worth a thousand words, what is a video worth? (2014), https://www.huffpost.com/entry/if-a-picture-video-production_b_4996655
2014
Earlier work this paper cites.
Hinton, G., Vinyals, O., Dean, J.: Distilling the knowledge in a neural network. Advances in Neural Information Processing Systems (NIPS) (2015)
2015
Earlier work this paper cites.
Gupta, S., Hoffman, J., Malik, J.: Cross modal distillation for supervision transfer. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2827–2836 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Ye, Y., Zhao, Z., Li, Y., Chen, L., Xiao, J., Zhuang, Y.: Video question answering via attribute-augmented attention network learning. In: Proceedings of the 40th International ACM SIGIR conference on Research and Development in Information Retrieval. pp. 829–832 (2017)
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Zhang, P., Zhuo, T., Li, X.: Learning to select key frames for video question answering. ACM Multimedia Conference (2018)
2018
Earlier work this paper cites.
Zhang, Y., Xiang, T., Hospedales, T.M., Lu, H.: Deep mutual learning. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4320–4328 (2018)
2018
Earlier work this paper cites.
Fan, C., Zhuo, T., Zhang, P., Li, X.: Heterogeneous memory enhanced multimodal attention model for video question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1999–2008 (2019)
2019
Earlier work this paper cites.
Mullapudi, R.T., Chen, S., Zhang, K., Ramanan, D., Fatahalian, K.: Online model distillation for efficient video inference. In: Proceedings of the IEEE/CVF International conference on computer vision. pp. 3573–3582 (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Jiang, J., Chen, Z., Lin, H., Zhao, X., Gao, Y.: Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 11101–11108 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Pan, B., Cai, H., Huang, D.A., Lee, K.H., Gaidon, A., Adeli, E., Niebles, J.C.: Spatio-temporal graph for video captioning with knowledge distillation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10870–10879 (2020)
2020
Earlier work this paper cites.
Amrani, E., Ben-Ari, R., Rotman, D., Bronstein, A.: Noise estimation using density estimation for self-supervised multimodal learning. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 35, pp. 6644–6652 (2021)
2021
Earlier work this paper cites.
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., Schmid, C.: Vivit: A video vision transformer. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 6836–6846 (2021)
2021
Earlier work this paper cites.
Bertasius, G., Wang, H., Torresani, L.: Is space-time attention all you need for video understanding? In: ICML. vol. 2, p. 4 (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Khani, M., Hamadanian, P., Nasr-Esfahany, A., Alizadeh, M.: Real-time video inference on edge devices via adaptive model streaming. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4572–4582 (2021)
2021
Earlier work this paper cites.
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., Hoi, S.C.H.: Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems 34
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PMLR (2021)
2021
Cited alongside, same era.
Wu, B., Yu, S., Chen, Z., Tenenbaum, J.B., Gan, C.: Star: A benchmark for situated reasoning in real-world videos. In: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) (2021)
2021
Cited alongside, same era.
Xiao, J., Shang, X., Yao, A., Chua, T.S.: Next-qa: Next phase of question-answering to explaining temporal actions. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 9777–9786 (2021)
2021
Cited alongside, same era.
Yang, A., Miech, A., Sivic, J., Laptev, I., Schmid, C.: Just ask: Learning to answer questions from millions of narrated videos. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 1686–1697 (2021)
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., Cao, Y.: Eva: Exploring the limits of masked visual representation learning at scale. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 19358–19369 (2023)
2023
Closest in time.
Gao, D., Zhou, L., Ji, L., Zhu, L., Yang, Y., Shou, M.Z.: Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 14773–14783 (2023)
2023
Closest in time.
Heigold, G., Minderer, M., Gritsenko, A., Bewley, A., Keysers, D., Lučić, M., Yu, F., Kipf, T.: Video owl-vit: Temporally-consistent open-world localization in video. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 13802–13811 (October 2023)
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J.S., Cao, J., Farhadi, A., Choi, Y.: Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems 34
2021
Cited alongside, same era.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O.K., Aggarwal, K., Som, S., Piao, S., Wei, F.: Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Li, J., Li, D., Xiong, C., Hoi, S.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning. pp. 12888–12900. PMLR (2022)
2022
Cited alongside, same era.
Liu, Y., Xiong, P., Xu, L., Cao, S., Jin, Q.: Ts2-net: Token shift and selection transformer for text-video retrieval. In: Proceedings of the European Conference on Computer Vision (ECCV) (2022)
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Wang, L., Qiao, Y.: Uniformerv2: Unlocking the potential of image vits for video understanding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 1632–1643 (October 2023)
2023
Closest in time.
Lin, Y., Wei, C., Wang, H., Yuille, A., Xie, C.: Smaug: Sparse masked autoencoder for efficient video-language pre-training. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 2459–2469 (2023)
2023
Closest in time.
2023
Closest in time.
Pramanick, S., Song, Y., Nag, S., Lin, K.Q., Shah, H., Shou, M.Z., Chellappa, R., Zhang, P.: Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 5285–5297 (2023)
2023
Closest in time.
Qing, Z., Zhang, S., Huang, Z., Zhang, Y., Gao, C., Zhao, D., Sang, N.: Disentangling spatial and temporal learning for efficient image-to-video transfer learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 13934–13944 (October 2023)
2023
Closest in time.
Shao, Z., Wan, J., Zong, L.: A video question answering model based on knowledge distillation. Information 14
2023
Closest in time.
Shen, Y., Wang, X., Gao, P., Lin, M.: Auxiliary modality learning with generalized curriculum distillation (2023)
2023
Closest in time.
Shen, Y., Yang, L., Wang, X., Lin, M.C.: Small-shot multi-modal distillation for vision-based autonomous steering. In: 2023 IEEE International Conference on Robotics and Automation (ICRA). pp. 7763–7770. IEEE (2023)
2023
Closest in time.
2023
Closest in time.
Wang, J., Ge, Y., Yan, R., Ge, Y., Lin, K.Q., Tsutsui, S., Lin, X., Cai, G., Wu, J., Shan, Y., et al.: All in one: Exploring unified video-language pre-training. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 6598–6608 (2023)
2023
Closest in time.
Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Yuan, L., Jiang, Y.G.: Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 6312–6322 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Wang, Z., Sung, Y.L., Cheng, F., Bertasius, G., Bansal, M.: Unified coarse-to-fine alignment for video-text retrieval. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2816–2827 (October 2023)
2023
Closest in time.
2023
Closest in time.
Xue, H., Sun, Y., Liu, B., Fu, J., Song, R., Li, H., Luo, J.: Clip-vip: Adapting pre-trained image-text model to video-language representation alignment (2023)
2023
Closest in time.
2023
Closest in time.