Fetching the paper…
Reading the bibliography…
Recent dominant methods for video-language pre-training (VLP) learn transferable representations from the raw pixels in an end-to-end manner to achieve advanced performance on downstream video-language retrieval.
Egly, R., Driver, J., Rafal, R.D.: Shifting visual attention between objects and locations: evidence from normal and parietal lesion subjects. p. 161 (1994)
1994
Earlier work this paper cites.
Scholl, B.J.: Objects and attention: The state of the art pp. 1–46 (2001)
2001
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. pp. 770–778 (2016)
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Hershey, S., Chaudhuri, S., Ellis, D.P., Gemmeke, J.F., Jansen, A., Moore, R.C., Plakal, M., Platt, D., Saurous, R.A., Seybold, B., et al.: Cnn architectures for large-scale audio classification. In: ICSAPP. pp. 131–135 (2017)
2017
Earlier work this paper cites.
Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: CVPR. pp. 4700–4708 (2017)
2017
Earlier work this paper cites.
Jang, Y., Song, Y., Yu, Y., Kim, Y., Kim, G.: Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In: CVPR. pp. 2758–2766 (2017)
2017
Earlier work this paper cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV pp. 32–73 (2017)
2017
Earlier work this paper cites.
Li, S., Xiao, T., Li, H., Yang, W., Wang, X.: Identity-aware textual-visual matching with latent co-attention. In: CVPR. pp. 1890–1899 (2017)
2017
Earlier work this paper cites.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: CVPR. pp. 6077–6086 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Gao, J., Ge, R., Chen, K., Nevatia, R.: Motion-appearance co-memory networks for video question answering. In: CVPR. pp. 6576–6585 (2018)
2018
Earlier work this paper cites.
Hara, K., Kataoka, H., Satoh, Y.: Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In: CVPR. pp. 6546–6555 (2018)
2018
Earlier work this paper cites.
Lee, K.H., Chen, X., Hua, G., Hu, H., He, X.: Stacked cross attention for image-text matching. In: ECCV. pp. 201–216 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Mithun, N.C., Li, J., Metze, F., Roy-Chowdhury, A.K.: Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In: ICMR. pp. 19–27 (2018)
2018
Earlier work this paper cites.
Yu, Y., Kim, J., Kim, G.: A joint sequence fusion model for video question answering and retrieval. In: ECCV. pp. 471–487 (2018)
2018
Earlier work this paper cites.
Zhang, B., Hu, H., Sha, F.: Cross-modal and hierarchical modeling of video and text. In: ECCV. pp. 374–390 (2018)
2018
Earlier work this paper cites.
Aafaq, N., Akhtar, N., Liu, W., Gilani, S.Z., Mian, A.: Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 12487–12496 (2019)
2019
Cited alongside, same era.
Fan, C., Zhang, X., Zhang, S., Wang, W., Zhang, C., Huang, H.: Heterogeneous memory enhanced multimodal attention model for video question answering. In: CVPR. pp. 1999–2007 (2019)
2019
Cited alongside, same era.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: ICCV. pp. 6202–6211 (2019)
2019
Cited alongside, same era.
Li, X., Song, J., Gao, L., Liu, X., Huang, W., He, X., Gan, C.: Beyond rnns: Positional self-attention with co-attention for video question answering. In: AAAI. pp. 8658–8665 (2019)
2019
Cited alongside, same era.
Pan, B., Cai, H., Huang, D.A., Lee, K.H., Gaidon, A., Adeli, E., Niebles, J.C.: Spatio-temporal graph for video captioning with knowledge distillation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10870–10879 (2020)
2020
Later among the works it cites.
Zheng, Q., Wang, C., Tao, D.: Syntax-aware action targeting for video captioning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 13096–13105 (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Amrani, E., Ben-Ari, R., Rotman, D., Bronstein, A.: Noise estimation using density estimation for self-supervised multimodal learning pp. 6644–6652 (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, C., Mao, Z., Liu, A.A., Zhang, T., Wang, B., Zhang, Y.: Focus your attention: A bidirectional focal attention network for image-text matching. In: ACMMM. pp. 3–11 (2019)
2019
Cited alongside, same era.
Liu, Y., Albanie, S., Nagrani, A., Zisserman, A.: Use what you have: Video retrieval using representations from collaborative experts. BMVC (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Sun, C., Myers, A., Vondrick, C., Murphy, K., Schmid, C.: Videobert: A joint model for video and language representation learning. In: ICCV. pp. 7464–7473 (2019)
2019
Cited alongside, same era.
Wang, Z., Liu, X., Li, H., Sheng, L., Yan, J., Wang, X., Shao, J.: Camp: Cross-modal adaptive message passing for text-image retrieval. In: ICCV. pp. 5764–5773 (2019)
2019
Cited alongside, same era.
Zhang, J., Peng, Y.: Object-aware aggregation with bidirectional temporal graph for video captioning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8327–8336 (2019)
2019
Cited alongside, same era.
Chen, S., Zhao, Y., Jin, Q., Wu, Q.: Fine-grained video-text retrieval with hierarchical graph reasoning. In: CVPR. pp. 10638–10647 (2020)
2020
Cited alongside, same era.
Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., Liu, J.: Uniter: Universal image-text representation learning. In: ECCV. pp. 104–120 (2020)
2020
Cited alongside, same era.
Bain, M., Nagrani, A., Varol, G., Zisserman, A.: Frozen in time: A joint video and image encoder for end-to-end retrieval. In: ICCV. pp. 1728–1738 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Changpinyo, S., Sharma, P., Ding, N., Soricut, R.: Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In: CVPR (2021)
2021
Later among the works it cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. ICLR (2021)
2021
Later among the works it cites.
Dzabraev, M., Kalashnikov, M., Komkov, S., Petiushko, A.: Mdmmt: Multidomain multimodal transformer for video retrieval. In: CVPR. pp. 3354–3363 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Lei, J., Li, L., Zhou, L., Gan, Z., Berg, T.L., Bansal, M., Liu, J.: Less is more: Clipbert for video-and-language learning via sparse sampling. In: CVPR. pp. 7331–7341 (2021)
2021
Later among the works it cites.
Patrick, M., Huang, P.Y., Asano, Y., Metze, F., Hauptmann, A.G., Henriques, J.F., Vedaldi, A.: Support-set bottlenecks for video-text representation learning. In: ICLR (2021)
2021
Later among the works it cites.
Rouditchenko, A., Boggust, A., Harwath, D., Thomas, S., Kuehne, H., Chen, B., Panda, R., Feris, R., Kingsbury, B., Picheny, M., et al.: Cascaded multilingual audio-visual learning from videos. Interspeech pp. 3006–3010 (2021)
2021
Later among the works it cites.
Wang, J., Bao, B., Xu, C.: Dualvgr: A dual-visual graph reasoning unit for video question answering. TMM (2021)
2021
Later among the works it cites.
Wang, W., Zhang, M., Chen, R., Cai, G., Zhou, P., Peng, P., Guo, X., Wu, J., Sun, X.: Dig into multi-modal cues for video retrieval with hierarchical alignment. In: IJCAI (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Yang, J., Bisk, Y., Gao, J.: Taco: Token-aware cascade contrastive learning for video-text alignment. In: ICCV. pp. 11562–11572 (2021)
2021
Later among the works it cites.
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J.S., Cao, J., Farhadi, A., Choi, Y.: Merlot: Multimodal neural script knowledge models. In: NIPS (2021)
2021
Later among the works it cites.
Ali, A., Schwartz, I., Hazan, T., Wolf, L.: Video and text matching with conditioned embeddings. In: CVPR. pp. 1565–1574 (2022)
2022
Closest in time.
Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., Xu, C.: FILIP: Fine-grained interactive language-image pre-training. In: ICLR (2022)
2022
Closest in time.