Fetching the paper…
Reading the bibliography…
Temporal action detection (TAD) aims to determine the semantic label and the temporal interval of every action instance in an untrimmed video.
F. Caba Heilbron, J. Carlos Niebles, and B. Ghanem, “Fast temporal activity proposals for efficient detection of human actions in untrimmed videos,” in CVPR , 2016, pp. 1914–1923
1923
Earlier work this paper cites.
S. Ma, L. Sigal, and S. Sclaroff, “Learning activity progression in lstms for activity detection and early detection,” in CVPR , June 2016, pp. 1942–1950
1950
Earlier work this paper cites.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “ActivityNet: A large-scale video benchmark for human activity understanding,” in CVPR , 2015, pp. 961–970
2015
Earlier work this paper cites.
Z. Shou, D. Wang, and S.-F. Chang, “Temporal action localization in untrimmed videos via multi-stage cnns,” in CVPR , 2016, pp. 1049–1058
2016
Earlier work this paper cites.
A. Richard and J. Gall, “Temporal action detection using a statistical language model,” in CVPR , 2016, pp. 3131–3140
2016
Earlier work this paper cites.
J. Yuan, B. Ni, X. Yang, and A. A. Kassim, “Temporal action localization with pyramid of score distribution features,” in CVPR , 2016, pp. 3093–3102
2016
Earlier work this paper cites.
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem, “Daps: Deep action proposals for action understanding,” in ECCV , 2016, pp. 768–784
2016
Earlier work this paper cites.
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei, “End-to-end learning of action detection from frame glimpses in videos,” in CVPR , 2016, pp. 2678–2687
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in ECCV , 2016, pp. 20–36
2016
Earlier work this paper cites.
H. Xu, A. Das, and K. Saenko, “R-c3d: region convolutional 3d network for temporal activity detection,” in ICCV , 2017, pp. 5794–5803
2017
Earlier work this paper cites.
T. Lin, X. Zhao, and Z. Shou, “Single shot temporal action detection,” in ACM MM , 2017, pp. 988–996
2017
Earlier work this paper cites.
Z.-H. Yuan, J. C. Stroud, T. Lu, and J. Deng, “Temporal action localization by structured maximal sums,” in CVPR , 2017, pp. 3684–3692
2017
Earlier work this paper cites.
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin, “Temporal action detection with structured segment networks,” ICCV , pp. 2914–2923, 2017
2017
Earlier work this paper cites.
C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks for action segmentation and detection,” in CVPR , 2017, pp. 156–165
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR , 2017, pp. 4724–4733
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
X. Dai, B. Singh, G. Zhang, L. S. Davis, and Y. Q. Chen, “Temporal context network for activity localization in videos,” in ICCV , 2017, pp. 5727–5736
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah, “The THUMOS challenge on action recognition for videos “in the wild”,” pp. 1–23, 2017
2017
Earlier work this paper cites.
F. C. Heilbron, W. Barrios, V. Escorcia, and B. Ghanem, “Scc: Semantic context cascade for efficient action detection.” in CVPR , 2017, pp. 3175–3184
2017
Earlier work this paper cites.
J. Gao, Z. Yang, C. Sun, K. Chen, and R. Nevatia, “Turn tap: Temporal unit regression network for temporal action proposals,” in ICCV , 2017, pp. 3648–3656
2017
Earlier work this paper cites.
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. C. Niebles, “Sst: Single-stream temporal action proposals,” in CVPR , 2017, pp. 6373–6382
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2017, pp. 1–18
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Z. Shou, J. Chan, A. Zareian, K. Miyazawa, and S.-F. Chang, “Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos,” in ICCV , 2017, pp. 1417–1426
2017
Cited alongside, same era.
Y.-W. Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar, “Rethinking the faster r-cnn architecture for temporal action localization,” in CVPR , 2018, pp. 1130–1139
2018
Cited alongside, same era.
H. Alwassel, F. Caba Heilbron, V. Escorcia, and B. Ghanem, “Diagnosing error in temporal action detectors,” in ECCV , 2018, pp. 256–272
2018
Cited alongside, same era.
F. Ma, L. Zhu, Y. Yang, S. Zha, G. Kundu, M. Feiszli, and Z. Shou, “Sf-net: Single-frame supervision for temporal action localization,” in ECCV . Springer, 2020, pp. 420–437
2020
Later among the works it cites.
L. Huang, Y. Huang, W. Ouyang, and L. Wang, “Relational prototypical network for weakly supervised temporal action localization,” in AAAI , vol. 34, no. 07, 2020, pp. 11 053–11 060
2020
Later among the works it cites.
Y. Zhai, L. Wang, W. Tang, Q. Zhang, J. Yuan, and G. Hua, “Two-stream consensus network for weakly-supervised temporal action localization,” in ECCV . Springer, 2020, pp. 37–54
2020
Later among the works it cites.
L. Zhu and Y. Yang, “Actbert: Learning global-local video-text representations,” in CVPR , 2020, pp. 8746–8755
2020
Later among the works it cites.
Y. Bai, Y. Wang, Y. Tong, Y. Yang, Q. Liu, and J. Liu, “Boundary content graph neural network for temporal action proposal generation,” in ECCV , 2020, pp. 121–137
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Lin, X. Zhao, H. Su, C. Wang, and M. Yang, “Bsn: Boundary sensitive network for temporal action proposal generation,” in ECCV , September 2018, pp. 3–21
2018
Cited alongside, same era.
J. Gao, K. Chen, and R. Nevatia, “Ctap: Complementary temporal action proposal generation,” in ECCV , September 2018, pp. 70–85
2018
Cited alongside, same era.
P. Nguyen, T. Liu, G. Prasad, and B. Han, “Weakly supervised action localization by sparse temporal pooling network,” in CVPR , 2018, pp. 6752–6761
2018
Cited alongside, same era.
S. Paul, S. Roy, and A. K. Roy-Chowdhury, “W-talc: Weakly-supervised temporal activity localization and classification,” in ECCV , September 2018, pp. 588–607
2018
Cited alongside, same era.
Z. Shou, H. Gao, L. Zhang, K. Miyazawa, and S.-F. Chang, “Autoloc: Weakly-supervised temporal action localization in untrimmed videos,” in ECCV , 2018, pp. 154–171
2018
Cited alongside, same era.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in CVPR , 2018, pp. 8739–8748
2018
Cited alongside, same era.
T. Lin, X. Liu, X. Li, E. Ding, and S. Wen, “Bmn: Boundary-matching network for temporal action proposal generation,” in ICCV , 2019, pp. 3889–3898
2019
Cited alongside, same era.
R. Zeng, W. Huang, M. Tan, Y. Rong, P. Zhao, J. Huang, and C. Gan, “Graph convolutional networks for temporal action localization,” in ICCV , 2019, pp. 7094–7103
2019
Cited alongside, same era.
2020
Later among the works it cites.
P. Zhao, L. Xie, C. Ju, Y. Zhang, Y. Wang, and Q. Tian, “Bottom-up temporal action localization with mutual regularization,” in ECCV , 2020
2020
Later among the works it cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in ICLR , 2021
2021
Closest in time.
X. Liu, Y. Hu, S. Bai, F. Ding, X. Bai, and P. H. S. Torr, “Multi-shot temporal event localization: A benchmark,” in CVPR , June 2021, pp. 12 596–12 606
2021
Closest in time.
C. Lin, C. Xu, D. Luo, Y. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and Y. Fu, “Learning salient boundary feature for anchor-free temporal action localization,” in CVPR , 2021, pp. 3320–3329
2021
Closest in time.
R. Su, D. Xu, L. Sheng, and W. Ouyang, “Pcg-tal: Progressive cross-granularity cooperation for temporal action localization,” IEEE Transactions on Image Processing , vol. 30, pp. 2103–2113, 2021
2021
Closest in time.
L. Huang, Y. Huang, W. Ouyang, and L. Wang, “Modeling sub-actions for weakly supervised temporal action localization,” IEEE Transactions on Image Processing , vol. 30, pp. 5154–5167, 2021
2021
Closest in time.
W. Yang, T. Zhang, Z. Mao, Y. Zhang, Q. Tian, and F. Wu, “Multi-scale structure-aware network for weakly supervised temporal action detection,” IEEE Transactions on Image Processing , vol. 30, pp. 5848–5861, 2021
2021
Closest in time.
A. Islam, C. Long, and R. Radke, “A hybrid attention mechanism for weakly-supervised temporal action localization,” in AAAI , vol. 35, no. 2, 2021, pp. 1637–1645
2021
Closest in time.
L. Huang, L. Wang, and H. Li, “Foreground-action consistency network for weakly supervised temporal action localization,” in ICCV , 2021, pp. 8002–8011
2021
Closest in time.
B. Tan, N. Xue, S. Bai, T. Wu, and G.-S. Xia, “Planetr: Structure-guided transformers for 3d plane recovery,” in CVPR , 2021, pp. 4186–4195
2021
Closest in time.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in ICML , July 2021, pp. 813–824
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Tan, J. Tang, L. Wang, and G. Wu, “Relaxed transformer decoders for direct action proposal generation,” in ICCV , October 2021, pp. 13 526–13 535
2021
Closest in time.
2021
Closest in time.
H. Alwassel, S. Giancola, and B. Ghanem, “Tsp: Temporally-sensitive pretraining of video encoders for localization tasks,” in ICCV Workshops , 2021, pp. 3166–3176
2021
Closest in time.
J.-N. Chen, S. Sun, J. He, P. H. Torr, A. Yuille, and S. Bai, “Transmix: Attend to mix for vision transformers,” in CVPR , 2022, pp. 12 135–12 144
2022
Closest in time.
D. Liang, X. Chen, W. Xu, Y. Zhou, and X. Bai, “Transcrowd: weakly-supervised crowd counting with transformers,” Science China Information Sciences , vol. 65, no. 6, pp. 1–14, 2022
2022
Closest in time.
X. Liu, S. Bai, and X. Bai, “An empirical study of end-to-end temporal action detection,” in CVPR , June 2022, pp. 20 010–20 019
2022
Closest in time.