Fetching the paper…
Reading the bibliography…
Long-form video understanding requires designing approaches that are able to temporally localize activities or language.
Rohrbach, M., Regneri, M., Andriluka, M., Amin, S., Pinkal, M., Schiele, B.: Script data for attribute-based recognition of composite activities. In: Proceedings of the European Conference on Computer Vision (ECCV) (2012)
2012
Earlier work this paper cites.
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., Pinkal, M.: Grounding Action Descriptions in Videos. ACL (2013)
2013
Earlier work this paper cites.
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: Large-scale video classification with convolutional neural networks. In: CVPR (2014)
2014
Earlier work this paper cites.
Heilbron, F.C., Escorcia, V., Ghanem, B., Niebles, J.C.: ActivityNet: A large-scale video benchmark for human activity understanding. In: CVPR (2015)
2015
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. In: ICLR (2015)
2015
Earlier work this paper cites.
Misra, I., Zitnick, C.L., Hebert, M.: Shuffle and learn: Unsupervised learning using temporal order verification. In: ECCV (2016)
2016
Earlier work this paper cites.
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hollywood in homes: Crowdsourcing data collection for activity understanding. In: ECCV (2016)
2016
Earlier work this paper cites.
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Van Gool, L.: Temporal segment networks: Towards good practices for deep action recognition. In: ECCV. Springer (2016)
2016
Earlier work this paper cites.
Anne Hendricks, L., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: Localizing Moments in Video With Natural Language. In: ICCV (2017)
2017
Earlier work this paper cites.
Buch, S., Escorcia, V., Shen, C., Ghanem, B., Carlos Niebles, J.: SST: single-stream temporal action proposals. In: CVPR (2017)
2017
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the Kinetics dataset. In: CVPR (2017)
2017
Earlier work this paper cites.
Gao, J., Sun, C., Yang, Z., Nevatia, R.: TALL: Temporal activity localization via language query. In: ICCV (2017)
2017
Earlier work this paper cites.
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A.: The Kinetics human action video dataset. arXiv preprint (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Lee, H., Huang, J., Singh, M., Yang, M.: Unsupervised representation learning by sorting sequences. In: ICCV (2017)
2017
Earlier work this paper cites.
Qiu, Z., Yao, T., Mei, T.: Learning spatio-temporal representation with pseudo-3d residual networks. In: ICCV (2017)
2017
Earlier work this paper cites.
Xu, H., Das, A., Saenko, K.: R-C3D: Region convolutional 3d network for temporal activity detection. In: ICCV (2017)
2017
Earlier work this paper cites.
Alwassel, H., Caba Heilbron, F., Escorcia, V., Ghanem, B.: Diagnosing error in temporal action detectors. In: ECCV (2018)
2018
Earlier work this paper cites.
Chen, J., Chen, X., Ma, L., Jie, Z., Chua, T.S.: Temporally Grounding Natural Sentence in Video. In: EMNLP (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Lin, T., Zhao, X., Su, H., Wang, C., Yang, M.: BSN: Boundary sensitive network for temporal action proposal generation. In: ECCV (2018)
2018
Earlier work this paper cites.
Liu, B., Yeung, S., Chou, E., Huang, D.A., Fei-Fei, L., Niebles, J.C.: Temporal Modular Networks for Retrieving Complex Compositional Activities in Videos. In: ECCV (2018)
2018
Earlier work this paper cites.
Liu, M., Wang, X., Nie, L., Tian, Q., Chen, B., Chua, T.S.: Cross-modal moment localization in videos. In: ACM MM (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Song, X., Han, Y.: VAL: Visual-attention action localizer. In: PCM (2018)
2018
Earlier work this paper cites.
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: A closer look at spatiotemporal convolutions for action recognition. In: CVPR (2018)
2018
Cited alongside, same era.
Wei, D., Lim, J., Zisserman, A., Freeman, W.T.: Learning and using the arrow of time. In: CVPR (2018)
2018
Cited alongside, same era.
Wu, A., Han, Y.: Multi-modal circulant fusion for video-to-language and backward. In: IJCAI (2018)
2018
Cited alongside, same era.
Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K.: Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In: Proceedings of the European Conference on Computer Vision (ECCV) (September 2018)
2018
Cited alongside, same era.
Yu, A.W., Dohan, D., Luong, M.T., Zhao, R., Chen, K., Norouzi, M., Le, Q.V.: QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension. In: ICLR (2018)
Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li and Stephen Gould: Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention. In: WACV (2020)
2020
Later among the works it cites.
Han, T., Xie, W., Zisserman, A.: Memory-augmented dense predictive coding for video representation learning. In: ECCV (2020)
2020
Later among the works it cites.
Jenni, S., Meishvili, G., Favaro, P.: Video representation learning by recognizing temporal transformations. In: ECCV (2020)
2020
Later among the works it cites.
Li, G., Duan, N., Fang, Y., Gong, M., Jiang, D.: Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. In: The Thirty-Fourth AAAI Conference on Artificial Intelligence. pp. 11336–11344. AAAI Press (2020)
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Zhou, B., Andonian, A., Oliva, A., Torralba, A.: Temporal relational reasoning in videos. In: ECCV (2018)
2018
Cited alongside, same era.
Ge, R., Gao, J., Chen, K., Nevatia, R.: MAC: Mining activity concepts for language-based temporal localization. In: WACV (2019)
2019
Cited alongside, same era.
Ghosh, S., Agarwal, A., Parekh, Z., Hauptmann, A.: ExCL: Extractive Clip Localization Using Natural Language Descriptions. In: ACL (2019)
2019
Cited alongside, same era.
He, K., Girshick, R., Dollár, P.: Rethinking ImageNet pre-training. In: CVPR (2019)
2019
Cited alongside, same era.
Lin, T., Liu, X., Li, X., Ding, E., Wen, S.: BMN: Boundary-matching network for temporal action proposal generation. In: ICCV (2019)
2019
Cited alongside, same era.
Long, F., Yao, T., Qiu, Z., Tian, X., Luo, J., Mei, T.: Gaussian temporal awareness networks for action localization. In: CVPR (2019)
2019
Cited alongside, same era.
Lu, C., Chen, L., Tan, C., Li, X., Xiao, J.: DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization. In: EMNLP-IJCNLP (2019)
2019
Cited alongside, same era.
2020
Later among the works it cites.
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, Xiaolin Li: Rethinking the Bottom-Up Framework for Query-based Video Localization. In: AAAI (2020)
2020
Later among the works it cites.
Miech, A., Alayrac, J.B., Smaira, L., Laptev, I., Sivic, J., Zisserman, A.: End-to-End Learning of Visual Representations from Uncurated Instructional Videos. In: CVPR (2020)
2020
Later among the works it cites.
Mun, J., Cho, M., Han, B.: Local-Global Video-Text Interactions for Temporal Grounding. In: CVPR (2020)
2020
Later among the works it cites.
Qian, R., Meng, T., Gong, B., Yang, M., Wang, H., Belongie, S.J., Cui, Y.: Spatiotemporal contrastive video representation learning. arxiv preprint (2020)
2020
Later among the works it cites.
Wang, J., Jiao, J., Liu, Y.H.: Self-supervised video representation learning by pace prediction. In: ECCV (2020)
2020
Later among the works it cites.
Wang, J., Ma, L., Jiang, W.: Temporally Grounding Language Queries in Videos by Contextual Boundary-aware Prediction. In: AAAI (2020)
2020
Later among the works it cites.
Xu, M., Zhao, C., Rojas, D.S., Thabet, A., Ghanem, B.: G-TAD: Sub-graph localization for temporal action detection. In: CVPR (2020)
2020
Later among the works it cites.
Yao, Y., Liu, C., Luo, D., Zhou, Y., Ye, Q.: Video playback rate perception for self-supervised spatio-temporal representation learning. In: CVPR (2020)
2020
Later among the works it cites.
Zeng, R., Xu, H., Huang, W., Chen, P., Tan, M., Gan, C.: Dense regression network for video grounding. In: CVPR (2020)
2020
Later among the works it cites.
Zhang, H., Sun, A., Jing, W., Zhou, J.T.: Span-based localizing network for natural language video localization. In: ACL (2020)
2020
Later among the works it cites.
Zhang, S., Peng, H., Fu, J., Luo, J.: Learning 2d temporal adjacent networks for moment localization with natural language. In: AAAI (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Alwassel, H., Giancola, S., Ghanem, B.: TSP: Temporally-sensitive pretraining of video encoders for localization tasks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops (2021)
2021
Later among the works it cites.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q.V., Sung, Y., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. ICML (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Nag, S., Zhu, X., Xiang, T.: Few-shot temporal action localization with query adaptive transformer. In: arXiv (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763 (2021)
2021
Later among the works it cites.
Xu, H., Ghosh, G., Huang, P.Y., Okhonko, D., Aghajanyan, A., Feichtenhofer, F.M.L.Z.C.: Videoclip: Contrastive pre-training for zero-shot video-text understanding. In: EMNLP (2021)
2021
Later among the works it cites.
Xu, M., Pérez-Rúa, J.M., Escorcia, V., Martínez, B., Zhu, X., Zhang, L., Ghanem, B., Xiang, T.: Boundary-sensitive pre-training for temporal localization in videos. In: ICCV. pp. 7220–7230 (2021)
2021
Later among the works it cites.