Fetching the paper…
Reading the bibliography…
Current methods for video activity localisation over time assume implicitly that activity temporal boundaries labelled for model training are determined and precise.
Rohrbach, M., Regneri, M., Andriluka, M., Amin, S., Pinkal, M., Schiele, B.: Script data for attribute-based recognition of composite activities. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 144–157. Springer (2012)
2012
Earlier work this paper cites.
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., Pinkal, M.: Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics 1
2013
Earlier work this paper cites.
Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word representation. In: Conference on Empirical Methods in Natural Language Processing (EMNLP). pp. 1532–1543 (2014), http://www.aclweb.org/anthology/D14-1162
2014
Earlier work this paper cites.
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: Proceedings of the International Conference on Learning Representations (ICLR) (2015)
2015
Earlier work this paper cites.
Heilbron, F.C., Escorcia, V., Ghanem, B., Niebles, J.C.: Activitynet: A large-scale video benchmark for human activity understanding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 961–970 (2015). https://doi.org/10.1109/CVPR.2015.7298698
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hollywood in homes: Crowdsourcing data collection for activity understanding. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 510–526. Springer (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Anne Hendricks, L., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: Localizing moments in video with natural language. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). pp. 5803–5812 (2017)
2017
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 6299–6308 (2017)
2017
Earlier work this paper cites.
Gao, J., Sun, C., Yang, Z., Nevatia, R.: Tall: Temporal activity localization via language query. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). pp. 5267–5275 (2017)
2017
Earlier work this paper cites.
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Niebles, J.C.: Dense-captioning events in videos. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV) (2017)
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS). pp. 5998–6008 (2017)
2017
Cited alongside, same era.
Zhao, Y., Xiong, Y., Wang, L., Wu, Z., Tang, X., Lin, D.: Temporal action detection with structured segment networks. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2914–2923 (2017)
2017
Cited alongside, same era.
Chen, J., Chen, X., Ma, L., Jie, Z., Chua, T.S.: Temporally grounding natural sentence in video. In: Conference on Empirical Methods in Natural Language Processing (EMNLP). pp. 162–171 (2018)
2018
Cited alongside, same era.
Wang, H., Zha, Z.J., Chen, X., Xiong, Z., Luo, J.: Dual path interaction network for video moment localization. In: Proceedings of the ACM International Conference on Multimedia (MM). pp. 4116–4124 (2020)
2020
Later among the works it cites.
Zeng, R., Xu, H., Huang, W., Chen, P., Tan, M., Gan, C.: Dense regression network for video grounding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10287–10296 (2020)
2020
Later among the works it cites.
Zhang, H., Sun, A., Jing, W., Zhou, J.T.: Span-based localizing network for natural language video localization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 6543–6554. Association for Computational Linguistics, Online (Jul 2020), https://www.aclweb.org/anthology/2020.acl-main.585
2020
Later among the works it cites.
Zhang, S., Peng, H., Fu, J., Luo, J.: Learning 2d temporal adjacent networks for moment localization with natural language. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). vol. 34, pp. 12870–12877 (2020)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu, A.W., Dohan, D., Le, Q., Luong, T., Zhao, R., Chen, K.: Fast and accurate reading comprehension by combining self-attention and convolution. In: Proceedings of the International Conference on Learning Representations (ICLR). vol. 2 (2018)
2018
Cited alongside, same era.
Ge, R., Gao, J., Chen, K., Nevatia, R.: Mac: Mining activity concepts for language-based temporal localization. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). pp. 245–253. IEEE (2019)
2019
Cited alongside, same era.
Ghosh, S., Agarwal, A., Parekh, Z., Hauptmann, A.: ExCL: Extractive Clip Localization Using Natural Language Descriptions. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). pp. 1984–1990. Association for Computational Linguistics, Minneapolis, Minnesota (Jun 2019), https://www.aclweb.org/anthology/N19-1198
2019
Cited alongside, same era.
Yuan, Y., Ma, L., Wang, J., Liu, W., Zhu, W.: Semantic conditioned dynamic modulation for temporal sentence grounding in videos. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS). pp. 534–544 (2019)
2019
Cited alongside, same era.
Zhang, S., Su, J., Luo, J.: Exploiting temporal relationships in video moment localization with natural language. In: Proceedings of the ACM International Conference on Multimedia (MM). pp. 1230–1238 (2019)
2019
Cited alongside, same era.
Mayu Otani, Yuta Nakahima, E.R., Heikkilä, J.: Uncovering hidden challenges in query-based video moment retrieval. In: Proceedings of the British Machine Vision Conference (BMVC) (2020)
2020
Cited alongside, same era.
Mun, J., Cho, M., Han, B.: Local-global video-text interactions for temporal grounding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10810–10819 (2020)
2020
Cited alongside, same era.
2020
Later among the works it cites.
Huang, J., Liu, Y., Gong, S., Jin, H.: Cross-sentence temporal and semantic relations in video activity localisation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 7199–7208 (2021)
2021
Later among the works it cites.
Li, K., Guo, D., Wang, M.: Proposal-free video grounding with contextual pyramid network. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). vol. 35, pp. 1902–1910 (2021)
2021
Later among the works it cites.
Nan, G., Qiao, R., Xiao, Y., Liu, J., Leng, S., Zhang, H., Lu, W.: Interventional video grounding with dual contrastive learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 2765–2775 (2021)
2021
Later among the works it cites.
Wang, H., Zha, Z.J., Li, L., Liu, D., Luo, J.: Structured multi-level interaction network for video moment localization via language query. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 7026–7035 (2021)
2021
Later among the works it cites.
Xiao, S., Chen, L., Zhang, S., Ji, W., Shao, J., Ye, L., Xiao, J.: Boundary proposal network for two-stage natural language video localization. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). vol. 35, pp. 2986–2994 (2021)
2021
Later among the works it cites.
Zhao, Y., Zhao, Z., Zhang, Z., Lin, Z.: Cascaded prediction network via segment tree for temporal video grounding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4197–4206 (2021)
2021
Later among the works it cites.
Zhou, H., Zhang, C., Luo, Y., Chen, Y., Hu, C.: Embracing uncertainty: Decoupling and de-bias for robust temporal grounding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 8445–8454 (2021)
2021
Later among the works it cites.