Fetching the paper…
Reading the bibliography…
Weakly supervised temporal action localization, which aims at temporally locating action instances in untrimmed videos using only video-level class labels during training, is an important yet challenging problem in video analysis.
Wang, H., Kläser, A., Schmid, C., Liu, C.L.: Action recognition by dense trajectories. In: CVPR. pp. 3169–3176. IEEE (2011)
2011
Earlier work this paper cites.
Uijlings, J.R., Van De Sande, K.E., Gevers, T., Smeulders, A.W.: Selective search for object recognition. In: IJCV. vol. 104, pp. 154–171. Springer (2013)
2013
Earlier work this paper cites.
Wang, H., Schmid, C.: Action recognition with improved trajectories. In: ICCV. pp. 3551–3558 (2013)
2013
Earlier work this paper cites.
Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR. pp. 580–587 (2014)
2014
Earlier work this paper cites.
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., Darrell, T.: Caffe: Convolutional architecture for fast feature embedding. In: Proceedings of the 22nd ACM international conference on Multimedia. pp. 675–678. ACM (2014)
2014
Earlier work this paper cites.
Jiang, Y., Liu, J., Zamir, A.R., Toderici, G., Laptev, I., Shah, M., Sukthankar, R.: Thumos challenge: Action recognition with a large number of classes. In: Computer Vision-ECCV workshop 2014 (2014)
2014
Earlier work this paper cites.
Oneata, D., Verbeek, J., Schmid, C.: The lear submission at thumos2014. THUMOS Action Recognition challenge (2014)
2014
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Advances in neural information processing systems. pp. 568–576 (2014)
2014
Earlier work this paper cites.
Zitnick, C.L., Dollár, P.: Edge boxes: Locating object proposals from edges. In: ECCV. pp. 391–405. Springer (2014)
2014
Earlier work this paper cites.
Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J.: Activitynet: A large-scale video benchmark for human activity understanding. In: CVPR. pp. 961–970 (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: ICCV. pp. 4489–4497 (2015)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Bilen, H., Vedaldi, A.: Weakly supervised deep detection networks. In: CVPR. pp. 2846–2854 (2016)
2016
Earlier work this paper cites.
Feichtenhofer, C., Pinz, A., Zisserman, A.: Convolutional two-stream network fusion for video action recognition. In: CVPR. pp. 1933–1941 (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. pp. 770–778 (2016)
2016
Cited alongside, same era.
Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: Unified, real-time object detection. In: CVPR. pp. 779–788 (2016)
2016
Cited alongside, same era.
Redmon, J., Farhadi, A.: Yolo9000: Better, faster, stronger. arXiv preprint arXiv:1612.08242 (2016)
2016
Cited alongside, same era.
Richard, A., Gall, J.: Temporal action detection using a statistical language model. In: CVPR. pp. 3131–3140 (2016)
2016
Cited alongside, same era.
Shou, Z., Wang, D., Chang, S.F.: Temporal action localization in untrimmed videos via multi-stage cnns. In: CVPR. pp. 1049–1058 (2016)
2016
Cited alongside, same era.
2017
Later among the works it cites.
Lin, T.Y., Dollár, P., Girshick, R.B., He, K., Hariharan, B., Belongie, S.J.: Feature pyramid networks for object detection. In: CVPR. vol. 1, p. 4 (2017)
2017
Later among the works it cites.
Shou, Z., Chan, J., Zareian, A., Miyazawa, K., Chang, S.F.: Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In: CVPR. pp. 1417–1426. IEEE (2017)
2017
Later among the works it cites.
Singh, K.K., Lee, Y.J.: Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization. In: ICCV (2017)
2017
Later among the works it cites.
Wang, L., Xiong, Y., Lin, D., Van Gool, L.: Untrimmednets for weakly supervised action recognition and detection. In: CVPR. vol. 2 (2017)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Van Gool, L.: Temporal segment networks: Towards good practices for deep action recognition. In: ECCV. pp. 20–36. Springer (2016)
2016
Cited alongside, same era.
Wang, R., Tao, D.: Uts at activitynet 2016. ActivityNet Large Scale Activity Recognition Challenge 2016 p. 8 (2016)
2016
Cited alongside, same era.
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: Learning deep features for discriminative localization. In: CVPR. pp. 2921–2929 (2016)
2016
Cited alongside, same era.
Bai, P., Tang, X., Wang, X., Liu, W.: Multiple instance detection network with online instance classifier refinement. In: CVPR. pp. 4322–4328 (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
He, K., Gkioxari, G., Dollar, P., Girshick, R.: Mask-rcnn. arXiv:1703.06870v2 (2017)
2017
Cited alongside, same era.
2017
Later among the works it cites.
Wei, Y., Feng, J., Liang, X., Cheng, M.M., Zhao, Y., Yan, S.: Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In: CVPR. vol. 1, p. 3 (2017)
2017
Later among the works it cites.
2017
Later among the works it cites.
Xu, H., Das, A., Saenko, K.: R-c3d: region convolutional 3d network for temporal activity detection. In: ICCV. pp. 5794–5803 (2017)
2017
Later among the works it cites.
Yuan, Z.H., Stroud, J.C., Lu, T., Deng, J.: Temporal action localization by structured maximal sums. In: CVPR. vol. 2, p. 7 (2017)
2017
Later among the works it cites.
Zhao, Y., Xiong, Y., Wang, L., Wu, Z., Tang, X., Lin, D.: Temporal action detection with structured segment networks. In: ICCV. vol. 2 (2017)
2017
Later among the works it cites.
Chao, Y.W., Vijayanarasimhan, S., Seybold, B., Ross, D.A., Deng, J., Sukthankar, R.: Rethinking the faster r-cnn architecture for temporal action localization. In: CVPR. pp. 1130–1139 (2018)
2018
Closest in time.
2018
Closest in time.
Su, H., Zhao, X., Lin, T., Fei, H.: Weakly supervised temporal action detection with shot-based temporal pooling network. In: ICONIP (2018)
2018
Closest in time.
Zhang, J., Bargal, S.A., Lin, Z., Brandt, J., Shen, X., Sclaroff, S.: Top-down neural attention by excitation backprop. In: IJCV. vol. 126, pp. 1084–1102. Springer (2018)
2018
Closest in time.