Fetching the paper…
Reading the bibliography…
In this paper we revisit feature fusion, an old-fashioned topic, in the new context of text-to-video retrieval.
Smeulders, A.W.M., Worring, M., Santini, S., Gupta, A., Jain, R.C.: Content-based image retrieval at the end of the early years. TPAMI 22
2000
Earlier work this paper cites.
Snoek, C.G.M., Worring, M.: Multimodal video indexing: A review of the state-of-the-art. Multimedia Tools Appl. 25
2005
Earlier work this paper cites.
Yilmaz, E., Aslam, J.A.: Estimating average precision with incomplete and imperfect judgments. In: CIKM (2006)
2006
Earlier work this paper cites.
Over, P., Awad, G., Smeaton, A.F., Foley, C., Lanagan, J.: Creating a web-scale video collection for research. In: The 1st Workshop on Web-scale Multimedia Corpus (2009)
2009
Earlier work this paper cites.
Atrey, P.K., Hossain, M.A., El Saddik, A., Kankanhalli, M.S.: Multimodal fusion for multimedia analysis: A survey. Multimedia Systems 16
2010
Earlier work this paper cites.
Chen, D., Dolan, W.: Collecting highly parallel data for paraphrase evaluation. In: CVPR (2011)
2011
Earlier work this paper cites.
Li, Y., Song, Y., Cao, L., Tetreault, J.R., Goldberg, L., Jaimes, A., Luo, J.: TGIF: A new dataset and benchmark on animated gif description. In: CVPR (2015)
2015
Earlier work this paper cites.
Tran, D., Bourdev, L.D., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: ICCV (2015)
2015
Earlier work this paper cites.
Joulin, A., van der Maaten, L., Jabri, A., Vasilache, N.: Learning visual features from large weakly supervised data. In: ECCV (2016)
2016
Earlier work this paper cites.
Le, D.D., Phan, S., Nguyen, V.T., Renoust, B., Nguyen, T.A., Hoang, V.N., Ngo, T.D., Tran, M.T., Watanabe, Y., Klinkigt, M., Hiroke, A., Duong, Duc A. Miyao, Y., Satoh, S.: NII-HITACHI-UIT at TRECVID 2016. In: TRECVID (2016)
2016
Earlier work this paper cites.
Liang, J., Chen, J., Huang, P., Li, X., Jiang, L., Lan, Z., Pan, P., Fan, H., Jin, Q., Sun, J., Chen, Y.: Informedia @ TRECVID 2016. In: TRECVID (2016)
2016
Earlier work this paper cites.
Markatopoulou, F., Moumtzidou, A., Galanopoulos, D., Mironidis, T., Kaltsa, V., Ioannidou, A., Symeonidis, S., Avgerinakis, K., Andreadis, S., Gialampoukidis, I., Vrochidis, S., Briassouli, A., Mezaris, V., Kompatsiaris, I., Patras, I.: ITI-CERTH participation in TRECVID 2016. In: TRECVID (2016)
2016
Earlier work this paper cites.
Xu, J., Mei, T., Yao, T., Rui, Y.: MSR-VTT: A large video description dataset for bridging video and language. In: CVPR (2016)
2016
Earlier work this paper cites.
Dauphin, Y.N., Fan, A., Auli, M., Grangier, D.: Language modeling with gated convolutional networks. In: ICML (2017)
2017
Earlier work this paper cites.
Nguyen, P.A., Li, Q., Cheng, Z.Q., Lu, Y.J., Zhang, H., Wu, X., Ngo, C.W.: Vireo @ TRECVID 2017: Video-to-text, ad-hoc video search and video hyperlinking. In: TRECVID (2017)
2017
Earlier work this paper cites.
Snoek, C.G., Li, X., Xu, C., Koelma, D.C.: University of Amsterdam and Renmin university at TRECVID 2017: Searching video, detecting events and describing video. In: TRECVID (2017)
2017
Earlier work this paper cites.
Ueki, K., Hirakawa, K., Kikuchi, K., Ogawa, T., Kobayashi, T.: Waseda_Meisei at TRECVID 2017: Ad-hoc video search. In: TRECVID (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Xue, X., Nie, F., Wang, S., Chang, X., Stantic, B., Yao, M.: Multi-view correlated feature learning by uncovering shared component. In: AAAI (2017)
2017
Earlier work this paper cites.
Baltrušaitis, T., Ahuja, C., Morency, L.P.: Multimodal machine learning: A survey and taxonomy. TPAMI 41
2018
Earlier work this paper cites.
Bastan, M., Shi, X., Gu, J., Heng, Z., Zhuo, C., Sng, D., Kot, A.: NTU ROSE lab at TRECVID 2018: Ad-hoc video search and video to text. In: TRECVID (2018)
2018
Earlier work this paper cites.
Dong, J., Li, X., Snoek, C.G.: Predicting visual features from text for image and video caption retrieval. TMM 20
2018
Earlier work this paper cites.
Faghri, F., Fleet, D.J., Kiros, J.R., Fidler, S.: VSE++: improving visual-semantic embeddings with hard negatives. In: BMVC (2018)
2018
Earlier work this paper cites.
Huang, P.Y., Liang, J., Vaibhav, V., Chang, X., Hauptmann, A.: Informedia@TRECVID 2018: Ad-hoc video search with discrete and continuous representations. In: TRECVID (2018)
2018
Cited alongside, same era.
Ilse, M., Tomczak, J., Welling, M.: Attention-based deep multiple instance learning. In: ICML (2018)
2018
Cited alongside, same era.
Li, X., Dong, J., Xu, C., Cao, J., Wang, X., Yang, G.: Renmin University of China and Zhejiang Gongshang University at TRECVID 2018: Deep Cross-Modal Embeddings for Video-Text Retrieval. In: TRECVID (2018)
2018
Cited alongside, same era.
Mahajan, D., Girshick, R.B., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., van der Maaten, L.: Exploring the limits of weakly supervised pretraining. In: ECCV (2018)
2018
Cited alongside, same era.
Miech, A., Laptev, I., Sivic, J.: Learning a text-video embedding from incomplete and heterogeneous data. arXiv (2018)
Feichtenhofer, C.: X3d: Expanding architectures for efficient video recognition. In: CVPR (2020)
2020
Later among the works it cites.
Gabeur, V., Sun, C., Alahari, K., Schmid, C.: Multi-modal transformer for video retrieval. In: ECCV (2020)
2020
Later among the works it cites.
Li, X., Wan, W., Zhou, Y., Zhao, J., Wei, Q., Rong, J., Zhou, P., Xu, L., Lang, L., Liu, Y., Niu, C., Ding, D., Jin, X.: Deep multiple instance learning with spatial attention for rop case classification, instance selection and abnormality localization. In: ICPR (2020)
2020
Later among the works it cites.
Li, X., Zhou, F., Chen, A.: Renmin University of China at TRECVID 2020: Sentence Encoder Assembly for Ad-hoc Video Search. In: TRECVID (2020)
2020
Later among the works it cites.
Liu, M., Chen, X., Zhang, Y., Li, Y., Rehg, J.M.: Attention distillation for learning video representations. In: BMVC (2020)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Mithun, N.C., Li, J., Metze, F., Roy-Chowdhury, A.K.: Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In: ICMR (2018)
2018
Cited alongside, same era.
Woo, S., Park, J., Lee, J., Kweon, I.S.: CBAM: Convolutional block attention module. In: ECCV (2018)
2018
Cited alongside, same era.
Yu, Y., Kim, J., Kim, G.: A joint sequence fusion model for video question answering and retrieval. In: ECCV (2018)
2018
Cited alongside, same era.
Berns, F., Rossetto, L., Schoeffmann, K., Beecks, C., Awad, G.: V3C1 dataset: An evaluation of content characteristics. In: ICMR (2019)
2019
Cited alongside, same era.
Devlin, J., Chang, M., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: NAACL-HLT (2019)
2019
Cited alongside, same era.
Ghadiyaram, D., Tran, D., Mahajan, D.: Large-scale weakly-supervised pre-training for video action recognition. In: CVPR (2019)
2019
Cited alongside, same era.
Li, H., Chen, J., Hu, R., Yu, M., Chen, H., Xu, Z.: Action recognition using visual attention with reinforcement learning. In: MMM (2019)
2019
Cited alongside, same era.
2020
Later among the works it cites.
Mettes, P., Koelma, D.C., Snoek, C.G.M.: Shuffled ImageNet banks for video event detection and search. TOMM 16
2020
Later among the works it cites.
Wu, J., Ngo, C.W.: Interpretable embedding for ad-hoc video search. In: ACMMM (2020)
2020
Later among the works it cites.
Wu, J., Nguyen, P.A., Ngo, C.W.: Vireo@ TRECVID 2020 ad-hoc video search. In: TRECVID (2020)
2020
Later among the works it cites.
Yang, X., Dong, J., Cao, Y., Wang, X., Wang, M., Chua, T.S.: Tree-augmented cross-modal encoding for complex-query video retrieval. In: SIGIR (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Amrani, E., Ben-Ari, R., Rotman, D., Bronstein, A.: Noise estimation using density estimation for self-supervised multimodal learning. In: AAAI (2021)
2021
Closest in time.
Bain, M., Nagrani, A., Varol, G., Zisserman, A.: Frozen in time: A joint video and image encoder for end-to-end retrieval. In: ICCV (2021)
2021
Closest in time.
Bertasius, G., Wang, H., Torresani, L.: Is space-time attention all you need for video understanding? In: ICML (2021)
2021
Closest in time.
Chen, A., Hu, F., Wang, Z., Zhou, F., Li, X.: What matters for ad-hoc video search? a large-scale evaluation on TRECVID. In: ICCV Workshop on ViRal (2021)
2021
Closest in time.
Croitoru, I., Bogolin, S.V., Leordeanu, M., Jin, H., Zisserman, A., Albanie, S., Liu, Y.: TEACHTEXT: Crossmodal generalized distillation for text-video retrieval. In: ICCV (2021)
2021
Closest in time.
Dong, J., Li, X., Xu, C., Yang, X., Yang, G., Wang, X.: Dual encoding for video retrieval by text. TPAMI (2021)
2021
Closest in time.
2021
Closest in time.
Li, X., Zhou, F., Xu, C., Ji, J., Yang, G.: SEA: Sentence encoder assembly for video retrieval by textual queries. TMM 23
2021
Closest in time.
2021
Closest in time.
Patrick, M., Huang, P., Asano, Y.M., Metze, F., Hauptmann, A.G., Henriques, J.F., Vedaldi, A.: Support-set bottlenecks for video-text representation learning. In: ICLR (2021)
2021
Closest in time.
Portillo-Quintero, J.A., Ortiz-Bayliss, J.C., Terashima-Marín, H.: A straightforward framework for video retrieval using CLIP. In: MCPR (2021)
2021
Closest in time.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML (2021)
2021
Closest in time.