Fetching the paper…
Reading the bibliography…
Temporal action detection (TAD) is extensively studied in the video understanding community by generally following the object detection pipeline in images.
TDN: temporal difference networks for efficient action recognition, in: CVPR, pp. 1895–1904
Wang, L., Tong, Z., Ji, B., Wu, G., 2021b · 1904
Earlier work this paper cites.
Efficient non-maximum suppression, in: 18th International Conference on Pattern Recognition (ICPR 2006), 20-24 August 2006, Hong Kong, China, IEEE Computer Society. pp. 850–855
Neubeck, A., Gool, L.V., 2006 · 2006
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Jiang, Y.G., Liu, J., Roshan Zamir, A., Toderici, G., Laptev, I., Shah, M., Sukthankar, R., 2014 · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos, in: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N.D., Weinberger, K.Q. (Eds.), NIPS, pp. 568–576
Simonyan, K., Zisserman, A., 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding, in: CVPR, pp. 961–970
Heilbron, F.C., Escorcia, V., Ghanem, B., Niebles, J.C., 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, in: Bach, F.R., Blei, D.M. (Eds.), Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, JMLR.org. pp. 448–456
Ioffe, S., Szegedy, C., 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks, in: NIPS
Ren, S., He, K., Girshick, R.B., Sun, J., 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks, in: ICCV, pp. 4489–4497
Tran, D., Bourdev, L.D., Fergus, R., Torresani, L., Paluri, M., 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, IEEE Computer Society. pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
SSD: single shot multibox detector, in: ECCV
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S.E., Fu, C., Berg, A.C., 2016 · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition, in: ECCV, pp. 20–36
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Gool, L.V., 2016 · 2016
Earlier work this paper cites.
End-to-end, single-stream temporal action detection in untrimmed videos, in: BMVC
Buch, S., Escorcia, V., Ghanem, B., Fei-Fei, L., Niebles, J.C., 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset, in: CVPR, pp. 4724–4733
Carreira, J., Zisserman, A., 2017 · 2017
Earlier work this paper cites.
TURN TAP: temporal unit regression network for temporal action proposals, in: ICCV, pp. 3648–3656
Gao, J., Yang, Z., Sun, C., Chen, K., Nevatia, R., 2017 · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A., 2017 · 2017
Earlier work this paper cites.
Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video, in: Proceedings of the IEEE International Conference on Computer Vision, pp. 2886–2894
Moltisanti, D., Wray, M., Mayol-Cuevas, W., Damen, D., 2017 · 2017
Earlier work this paper cites.
Inception single shot multibox detector for object detection, in: 2017 IEEE International Conference on Multimedia & Expo Workshops, ICME Workshops, Hong Kong, China, July 10-14, 2017, IEEE Computer Society. pp. 549–554
Ning, C., Zhou, H., Song, Y., Tang, J., 2017 · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks, in: ICCV, pp. 5534–5542
Qiu, Z., Yao, T., Mei, T., 2017 · 2017
Earlier work this paper cites.
Convnet architecture search for spatiotemporal feature learning
Tran, D., Ray, J., Shou, Z., Chang, S., Paluri, M., 2017 · 2017
Earlier work this paper cites.
Attention is all you need, in: Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., Garnett, R. (Eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5998–6008
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
Untrimmednets for weakly supervised action recognition and detection, in: CVPR, pp. 6402–6411
Wang, L., Xiong, Y., Lin, D., Gool, L.V., 2017 · 2017
Cited alongside, same era.
R-C3D: region convolutional 3d network for temporal activity detection, in: ICCV
Xu, H., Das, A., Saenko, K., 2017 · 2017
Cited alongside, same era.
Diagnosing error in temporal action detectors, in: ECCV, Springer. pp. 264–280
Alwassel, H., Heilbron, F.C., Escorcia, V., Ghanem, B., 2018 · 2018
Cited alongside, same era.
Rethinking the faster R-CNN architecture for temporal action localization, in: CVPR, pp. 1130–1139
Chao, Y., Vijayanarasimhan, S., Seybold, B., Ross, D.A., Deng, J., Sukthankar, R., 2018 · 2018
Cited alongside, same era.
BSN: boundary sensitive network for temporal action proposal generation, in: Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y. (Eds.), ECCV, pp. 3–21
Lin, T., Zhao, X., Su, H., Wang, C., Yang, M., 2018 · 2018
Cited alongside, same era.
Something-else: Compositional action recognition with spatial-temporal interaction networks, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, Computer Vision Foundation / IEEE. pp. 1046–1056
Materzynska, J., Xiao, T., Herzig, R., Xu, H., Wang, X., Darrell, T., 2020 · 2020
Later among the works it cites.
G-TAD: sub-graph localization for temporal action detection, in: CVPR, Computer Vision Foundation / IEEE. pp. 10153–10162
Xu, M., Zhao, C., Rojas, D.S., Thabet, A.K., Ghanem, B., 2020 · 2020
Later among the works it cites.
Revisiting anchor mechanisms for temporal action localization
Yang, L., Peng, H., Zhang, D., Fu, J., Han, J., 2020 · 2020
Later among the works it cites.
Temporal action detection with structured segment networks
Zhao, Y., Xiong, Y., Wang, L., Wu, Z., Tang, X., Lin, D., 2020 · 2020
Later among the works it cites.
Distance-iou loss: Faster and better learning for bounding box regression, in: AAAI, pp. 12993–13000
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A closer look at spatiotemporal convolutions for action recognition, in: CVPR, pp. 6450–6459
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M., 2018 · 2018
Cited alongside, same era.
Non-local neural networks, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, Computer Vision Foundation / IEEE Computer Society. pp. 7794–7803
Wang, X., Girshick, R.B., Gupta, A., He, K., 2018b · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification, in: Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y. (Eds.), ECCV, pp. 318–335
Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K., 2018 · 2018
Cited alongside, same era.
Slowfast networks for video recognition, in: ICCV, pp. 6201–6210
Feichtenhofer, C., Fan, H., Malik, J., He, K., 2019 · 2019
Cited alongside, same era.
Video imprint segmentation for temporal action detection in untrimmed videos, in: The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, AAAI Press. pp. 8328–8335
Gao, Z., Wang, L., Zhang, Q., Niu, Z., Zheng, N., Hua, G., 2019 · 2019
Cited alongside, same era.
Multi-granularity generator for temporal action proposal, in: CVPR, pp. 3604–3613
Liu, Y., Ma, L., Zhang, Y., Liu, W., Chang, S., 2019 · 2019
Cited alongside, same era.
Gaussian temporal awareness networks for action localization, in: CVPR, pp. 344–353
Long, F., Yao, T., Qiu, Z., Tian, X., Luo, J., Mei, T., 2019 · 2019
Cited alongside, same era.
Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., Ren, D., 2020 · 2020
Later among the works it cites.
Vivit: A video vision transformer, in: ICCV, pp. 6816–6826
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lucic, M., Schmid, C., 2021 · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?, in: Meila, M., Zhang, T. (Eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, PMLR. pp. 813–824
Bertasius, G., Wang, H., Torresani, L., 2021 · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, in: 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, OpenReview.net
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N., 2021 · 2021
Later among the works it cites.
Learning salient boundary feature for anchor-free temporal action localization, in: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, Computer Vision Foundation / IEEE. pp. 3320–3329
Lin, C., Xu, C., Luo, D., Wang, Y., Tai, Y., Wang, C., Li, J., Huang, F., Fu, Y., 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows, in: 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, IEEE. pp. 9992–10002
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B., 2021b · 2021
Later among the works it cites.
Neimark, D., Bar, O., Zohar, M., Asselmann, D., 2021 · 2021
Later among the works it cites.
Temporal context aggregation network for temporal action proposal refinement, in: CVPR, pp. 485–494
Qing, Z., Su, H., Gan, W., Wang, D., Wu, W., Wang, X., Qiao, Y., Yan, J., Gao, C., Sang, N., 2021 · 2021
Later among the works it cites.
BSN++: complementary boundary regressor with scale-balanced relation modeling for temporal action proposal generation, in: AAAI, pp. 2602–2610
Su, H., Gan, W., Wu, W., Qiao, Y., Yan, J., 2021 · 2021
Later among the works it cites.
Relaxed transformer decoders for direct action proposal generation, in: ICCV, pp. 13526–13535
Tan, J., Tang, J., Wang, L., Wu, G., 2021 · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention, in: Meila, M., Zhang, T. (Eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, PMLR. pp. 10347–10357
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H., 2021 · 2021
Later among the works it cites.
Towards high-quality temporal action detection with sparse proposals
Wu, J., Sun, P., Chen, S., Yang, J., Qi, Z., Ma, L., Luo, P., 2021 · 2021
Later among the works it cites.
DCAN: improving temporal action detection via dual context aggregation, in: AAAI, AAAI Press. pp. 248–257
Chen, G., Zheng, Y., Wang, L., Lu, T., 2022 · 2022
Closest in time.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., Wang, L., 2022 · 2022
Closest in time.
Actionformer: Localizing moments of actions with transformers
Zhang, C., Wu, J., Li, Y., 2022 · 2022
Closest in time.
VideoMAE V2: Scaling video masked autoencoders with dual masking, in: CVPR
Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., Qiao, Y., 2023 · 2023
Closest in time.