Fetching the paper…
Reading the bibliography…
Online action detection aims at the accurate action prediction of the current frame based on long historical observations.
M. Gao, Y. Zhou, R. Xu, R. Socher, and C. Xiong, “Woad: Weakly supervised online action detection in untrimmed videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 1915–1923
1923
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele, “A database for fine grained activity detection of cooking activities,” computer vision and pattern recognition , 2012
2012
Earlier work this paper cites.
K. Tang, L. Fei-Fei, and D. Koller, “Learning latent temporal structure for complex event detection,” computer vision and pattern recognition , 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
S. Karaman, L. Seidenari, and A. Del Bimbo, “Fast saliency based pooling of fisher encoded dense trajectories,” in ECCV THUMOS Workshop , vol. 1, no. 2, 2014, p. 5
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
H. Kuehne, J. Gall, and T. Serre, “An end-to-end generative framework for video segmentation and recognition,” workshop on applications of computer vision , 2015
2015
Earlier work this paper cites.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proceedings of the ieee conference on computer vision and pattern recognition , 2015, pp. 961–970
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on machine learning . PMLR, 2015, pp. 448–456
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
R. D. Geest, E. Gavves, A. Ghodrati, Z. Li, C. Snoek, and T. Tuytelaars, “Online action detection,” in European Conference on Computer Vision . Springer, 2016, pp. 269–284
2016
Earlier work this paper cites.
B. Singh, T. K. Marks, M. Jones, O. Tuzel, and M. Shao, “A multi-stream bi-directional recurrent neural network for fine-grained action detection,” computer vision and pattern recognition , 2016
2016
Earlier work this paper cites.
C. Lea, A. Reiter, R. Vidal, and G. D. Hager, “Segmental spatiotemporal cnns for fine-grained action segmentation,” european conference on computer vision , 2016
2016
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” arXiv: Sound , 2016
2016
Earlier work this paper cites.
C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks for action segmentation and detection,” computer vision and pattern recognition , 2016
2016
Earlier work this paper cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” european conference on computer vision , 2016
2016
Earlier work this paper cites.
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem, “Daps: Deep action proposals for action understanding,” european conference on computer vision , 2016
2016
Earlier work this paper cites.
R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “Netvlad: Cnn architecture for weakly supervised place recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 5297–5307
2016
Earlier work this paper cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in European conference on computer vision . Springer, 2016, pp. 20–36
2016
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah, “The thumos challenge on action recognition for videos “in the wild”,” Computer Vision and Image Understanding , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Lin, X. Zhao, and Z. Shou, “Single shot temporal action detection,” acm multimedia , 2017
2017
Earlier work this paper cites.
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. C. Niebles, “Sst: Single-stream temporal action proposals,” computer vision and pattern recognition , 2017
2017
Earlier work this paper cites.
S. Buch, V. Escorcia, B. Ghanem, L. Fei-Fei, and J. C. Niebles, “End-to-end, single-stream temporal action detection in untrimmed videos.” british machine vision conference , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah, “The thumos challenge on action recognition for videos “in the wild”,” Computer Vision and Image Understanding , vol. 155, pp. 1–23, 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6299–6308
2017
Cited alongside, same era.
Z. Shou, J. Chan, A. Zareian, K. Miyazawa, and S.-F. Chang, “Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 5734–5743
2017
Cited alongside, same era.
2017
Cited alongside, same era.
V. Ramanishka, Y.-T. Chen, T. Misu, and K. Saenko, “Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2020
Later among the works it cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International Conference on Machine Learning . PMLR, 2020, pp. 5156–5165
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Xu, Y. Xiong, H. Chen, X. Li, W. Xia, Z. Tu, and S. Soatto, “Long short-term transformer for online action detection,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
R. De Geest and T. Tuytelaars, “Modeling temporal structure with lstm for online action detection,” in 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 2018, pp. 1549–1557
2018
Cited alongside, same era.
P. Lei and S. Todorovic, “Temporal deformable residual networks for action segmentation in videos,” computer vision and pattern recognition , 2018
2018
Cited alongside, same era.
Y. Liu, L. Ma, Y. Zhang, W. Liu, and S.-F. Chang, “Multi-granularity generator for temporal action proposal,” computer vision and pattern recognition , 2018
2018
Cited alongside, same era.
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” in International conference on machine learning . PMLR, 2018, pp. 4055–4064
2018
Cited alongside, same era.
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick, “Long-term feature banks for detailed video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 284–293
2019
Cited alongside, same era.
Y. A. Farha and J. Gall, “Ms-tcn: Multi-stage temporal convolutional network for action segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3575–3584
2019
Cited alongside, same era.
M. Xu, M. Gao, Y.-T. Chen, L. S. Davis, and D. J. Crandall, “Temporal recurrent networks for online action detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 5532–5541
2019
Cited alongside, same era.
T. Lin, X. Liu, L. Xin, E. Ding, and S. Wen, “Bmn: Boundary-matching network for temporal action proposal generation,” international conference on computer vision , 2019
2019
Cited alongside, same era.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 4651–4664
2021
Later among the works it cites.
X. Wang, S. Zhang, Z. Qing, Y. Shao, Z. Zuo, C. Gao, and N. Sang, “Oadtr: Online action detection with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 7565–7575
2021
Later among the works it cites.
C. Lin, C. Xu, D. Luo, Y. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and Y. Fu, “Learning salient boundary feature for anchor-free temporal action localization,” computer vision and pattern recognition , 2021
2021
Later among the works it cites.
Z. Qing, H. Su, W. Gan, D. Wang, W. Wu, X. Wang, Y. Qiao, J. Yan, C. Gao, and N. Sang, “Temporal context aggregation network for temporal action proposal refinement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 485–494
2021
Later among the works it cites.
Z. Zhu, W. Tang, L. Wang, N. Zheng, and G. Hua, “Enriching local and global contexts for temporal action localization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 516–13 525
2021
Later among the works it cites.
J. Tan, J. Tang, L. Wang, and G. Wu, “Relaxed transformer decoders for direct action proposal generation,” international conference on computer vision , 2021
2021
Later among the works it cites.
D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 3163–3172
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Li, Y. Zhang, C. Liu, B. Shuai, Y. Zhu, B. Brattoli, H. Chen, I. Marsic, and J. Tighe, “Vidtr: Video transformer without convolutions,” arXiv e-prints , pp. arXiv–2104, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “Vivit: A video vision transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6836–6846
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Tan, J. Tang, L. Wang, and G. Wu, “Relaxed transformer decoders for direct action proposal generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 526–13 535
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh, “Nyströmformer: A nyström-based algorithm for approximating self-attention,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 16, 2021, pp. 14 138–14 148
2021
Later among the works it cites.
Y. Tay, D. Bahri, D. Metzler, D.-C. Juan, Z. Zhao, and C. Zheng, “Synthesizer: Rethinking self-attention for transformer models,” in International conference on machine learning . PMLR, 2021, pp. 10 183–10 192
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. H. Kim, S. Nam, and S. J. Kim, “Temporally smooth online action detection using cycle-consistent future anticipation,” Pattern Recognition , 2021
2021
Later among the works it cites.
H. Eun, J. Moon, J. Park, C. Jung, and C. Kim, “Temporal filtering networks for online action detection,” Pattern Recognition , vol. 111, p. 107695, 2021
2021
Later among the works it cites.
J. Chen, G. Mittal, Y. Yu, Y. Kong, and M. Chen, “Gatehub: Gated history unit with background suppression for online action detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 19 925–19 934
2022
Closest in time.
L. Yang, J. Han, and D. Zhang, “Colar: Effective and efficient online action detection by consulting exemplars,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3160–3169
2022
Closest in time.
F. Yi, H. Wen, and T. Jiang, “Asformer: Transformer for action segmentation,” 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.