Fetching the paper…
Reading the bibliography…
Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes.
In: 2009 Twelfth IEEE Intl. workshop on performance evaluation of tracking and surveillance, pp. 1–6. IEEE (2009)
Ferryman, J., Shahrokni, A.: Pets2009: Dataset and challenge · 2009
Earlier work this paper cites.
In: ECCV (2010)
Eichner, M., Ferrari, V.: We are family: Joint pose estimation of multiple persons · 2010
Earlier work this paper cites.
In: bmvc (2010)
Johnson, S., Everingham, M.: Clustered pose and nonlinear appearance models for human pose estimation · 2010
Earlier work this paper cites.
In: 2011 Intl. Conf. on Computer Vision. IEEE (2011)
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: Hmdb: a large video database for human motion recognition · 2011
Earlier work this paper cites.
In: 2012 CVPR. IEEE (2012)
Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? the kitti vision benchmark suite · 2012
Earlier work this paper cites.
Soomro, K., Zamir, A.R., Shah, M.: Ucf101: A dataset of 101 human actions classes from videos in the wild · 2012
Earlier work this paper cites.
In: IEEE Intl. Conf. on Computer Vision, pp. 2720–2727 (2013)
Lu, C., Shi, J., Jia, J.: Abnormal event detection at 150 fps in matlab · 2013
Earlier work this paper cites.
ACM Trans. OMM 9
Mei, T., Tang, L.X., Tang, J., Hua, X.S.: Near-lossless semantic video summarization and its applications to video analysis · 2013
Earlier work this paper cites.
In: CVPR, pp. 3674–3681 (2013)
Sapp, B., Taskar, B.: Modec: Multimodal decomposable models for human pose estimation · 2013
Earlier work this paper cites.
In: CVPR (2014)
Andriluka, M., Pishchulin, L., Gehler, P., Schiele, B.: 2d human pose estimation: New benchmark and state of the art analysis · 2014
Earlier work this paper cites.
In: ECCV (2014)
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context · 2014
Earlier work this paper cites.
In: Advances in neural information processing systems, pp. 91–99 (2015)
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks · 2015
Earlier work this paper cites.
In: IEEE Intl. Conf. on Computer Vision, pp. 4041–4049 (2015)
Veeriah, V., Zhuang, N., Qi, G.J.: Differential recurrent neural networks for action recognition · 2015
Earlier work this paper cites.
In: CVPR, pp. 1293–1301 (2015)
Xiaohan Nie, B., Xiong, C., Zhu, S.C.: Joint action recognition and pose estimation from video · 2015
Earlier work this paper cites.
In: 2016 IEEE Intl. Conf. on Image Processing (ICIP), pp. 3464–3468. IEEE (2016)
Bewley, A., Ge, Z., Ott, L., Ramos, F., Upcroft, B.: Simple online and realtime tracking · 2016
Earlier work this paper cites.
IEEE Transactions on Image Processing 25
Du, Y., Fu, Y., Wang, L.: Representation learning of temporal dynamics for skeleton-based action recognition · 2016
Earlier work this paper cites.
In: CVPR, pp. 770–778 (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Earlier work this paper cites.
Milan, A., Leal-Taixé, L., Reid, I., Roth, S., Schindler, K.: Mot16: A benchmark for multi-object tracking · 2016
Earlier work this paper cites.
In: CVPR, pp. 4929–4937 (2016)
Pishchulin, L., Insafutdinov, E., Tang, S., Andres, B., Andriluka, M., Gehler, P.V., Schiele, B.: Deepcut: Joint subset partition and labeling for multi person pose estimation · 2016
Earlier work this paper cites.
In: ECCV, pp. 17–35 (2016)
Ristani, E., Solera, F., Zou, R., Cucchiara, R., Tomasi, C.: Performance measures and a data set for multi-target, multi-camera tracking · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1010–1019 (2016)
Shahroudy, A., Liu, J., Ng, T.T., Wang, G.: Ntu rgb+ d: A large scale dataset for 3d human activity analysis · 2016
Earlier work this paper cites.
In: 2017 14th IEEE Intl. Conf. on Advanced Video and Signal Based Surveillance (AVSS). IEEE (2017)
Bochinski, E., Eiselein, V., Sikora, T.: High-speed tracking-by-detection without using image information · 2017
Earlier work this paper cites.
In: CVPR, pp. 6299–6308 (2017)
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset · 2017
Earlier work this paper cites.
In: IEEE Intl. Conf. on Computer Vision (2017)
Fang, H.S., Xie, S., Tai, Y.W., Lu, C.: Rmpe: Regional multi-person pose estimation · 2017
Cited alongside, same era.
In: Proceedings of the IEEE international conference on computer vision, pp. 5842–5850 (2017)
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al.: The” something something” video database for learning and evaluating visual common sense · 2017
Cited alongside, same era.
In: 2017 12th IEEE Intl. Conf. on Automatic Face & Gesture Recognition (FG 2017), pp. 438–445. IEEE (2017)
Iqbal, U., Garbade, M., Gall, J.: Pose for action-action for pose · 2017
Cited alongside, same era.
In: Proceedings of the IEEE International Conference on Computer Vision, pp. 4405–4413 (2017)
Kalogeiton, V., Weinzaepfel, P., Ferrari, V., Schmid, C.: Action tubelet detector for spatio-temporal action localization · 2017
Cited alongside, same era.
IEEE Transactions on Image Processing 27
Liu, J., Wang, G., Duan, L.Y., Abdiyeva, K., Kot, A.C.: Skeleton-based human action recognition with global context-aware attention lstm networks · 2017
In: Proceedings of the AAAI conference on artificial intelligence, vol. 32 (2018)
Yan, S., Xiong, Y., Lin, D.: Spatial temporal graph convolutional networks for skeleton-based action recognition · 2018
Later among the works it cites.
In: IEEE Intl. Conf. on computer vision (2019)
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition · 2019
Later among the works it cites.
In: CVPR, pp. 244–253 (2019)
Girdhar, R., Carreira, J., Doersch, C., Zisserman, A.: Video action transformer network · 2019
Later among the works it cites.
In: CVPR, pp. 10863–10872 (2019)
Li, J., Wang, C., Zhu, H., Mao, Y., Fang, H.S., Lu, C.: Crowdpose: Efficient crowded scenes pose estimation and a new benchmark · 2019
Later among the works it cites.
Ning, G., Huang, H.: Lighttrack: A generic framework for online top-down human pose tracking · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: Proceedings of the IEEE International Conference on Computer Vision, pp. 3637–3646 (2017)
Singh, G., Saha, S., Sapienza, M., Torr, P.H., Cuzzolin, F.: Online real-time multiple spatiotemporal action localisation and prediction · 2017
Cited alongside, same era.
In: Advances in neural information processing systems, pp. 5998–6008 (2017)
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need · 2017
Cited alongside, same era.
In: 2017 IEEE Intl. Conf. on image processing, pp. 3645–3649. IEEE (2017)
Wojke, N., Bewley, A., Paulus, D.: Simple online and realtime tracking with a deep association metric · 2017
Cited alongside, same era.
In: CVPR, pp. 5167–5176 (2018)
Andriluka, M., Iqbal, U., Insafutdinov, E., Pishchulin, L., Milan, A., Gall, J., Schiele, B.: Posetrack: A benchmark for human pose estimation and tracking · 2018
Cited alongside, same era.
In: ICME (2018)
Chen, L., Ai, H., Zhuang, Z., Shang, C.: Real-time multiple people tracking with deeply learned candidate selection and person re-identification · 2018
Cited alongside, same era.
Girdhar, R., Carreira, J., Doersch, C., Zisserman, A.: A better baseline for ava · 2018
Cited alongside, same era.
In: CVPR (2018)
Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., et al.: Ava: A video dataset of spatio-temporally localized atomic visual actions · 2018
Cited alongside, same era.
IEEE transactions on pattern analysis and machine intelligence (2019)
Shu, X., Tang, J., Qi, G., Liu, W., Yang, J.: Hierarchical long short-term concurrent memory for human interaction recognition · 2019
Later among the works it cites.
In: CVPR (2019)
Sun, K., Xiao, B., Liu, D., Wang, J.: Deep high-resolution representation learning for human pose estimation · 2019
Later among the works it cites.
Wang, Z., Zheng, L., Liu, Y., Wang, S.: Towards real-time multi-object tracking · 2019
Later among the works it cites.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 284–293 (2019)
Wu, C.Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., Girshick, R.: Long-term feature banks for detailed video understanding · 2019
Later among the works it cites.
In: European Conference on Computer Vision, pp. 455–472. Springer (2020)
Cai, Y., Wang, Z., Luo, Z., Yin, B., Du, A., Wang, H., Zhang, X., Zhou, X., Zhou, E., Sun, J.: Learning delicate local representations for multi-person pose estimation · 2020
Closest in time.
In: CVPR (2020)
Cheng, B., Xiao, B., Wang, J., Shi, H., Huang, T.S., Zhang, L.: Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation · 2020
Closest in time.
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 43
Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: The epic-kitchens dataset: Collection, challenges and baselines · 2020
Closest in time.
Dendorfer, P., Rezatofighi, H., Milan, A., Shi, J., Cremers, D., Reid, I., Roth, S., Schindler, K., Leal-Taixé, L.: Mot20: A benchmark for multi object tracking in crowded scenes · 2020
Closest in time.
Pattern Recognition 107
Peng, J., Wang, T., Lin, W., Wang, J., See, J., Wen, S., Ding, E.: Tpm: Multiple object tracking with tracklet-plane matching · 2020
Closest in time.
Zhang, Y., Wang, C., Wang, X., Zeng, W., Liu, W.: Fairmot: On the fairness of detection and re-identification in multiple object tracking · 2020
Closest in time.
In: European Conference on Computer Vision, pp. 474–490 (2020)
Zhou, X., Koltun, V., Krähenbühl, P.: Tracking objects as points · 2020
Closest in time.
In: ICML, vol. 2, p. 4 (2021)
Bertasius, G., Wang, H., Torresani, L.: Is space-time attention all you need for video understanding? · 2021
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14676–14686 (2021)
Geng, Z., Sun, K., Xiao, B., Zhang, Z., Wang, J.: Bottom-up human pose estimation via disentangled keypoint regression · 2021
Closest in time.
In: Proceedings of the 29th ACM International Conference on Multimedia, pp. 2158–2166 (2021)
Li, Y., Zhang, B., Li, J., Wang, Y., Lin, W., Wang, C., Li, J., Huang, F.: Lstc: Boosting atomic action detection with long-short-term context · 2021
Closest in time.
International journal of computer vision 129
Luiten, J., Osep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L., Leibe, B.: Hota: A higher order metric for evaluating multi-object tracking · 2021
Closest in time.
Advances in Neural Information Processing Systems (2021)
Yuan, Y., Fu, R., Huang, L., Lin, W., Zhang, C., Chen, X., Wang, J.: Hrformer: High-resolution vision transformer for dense predict · 2021
Closest in time.
IEEE Transactions on Multimedia (2022)
Chen, Y., Zhao, P., Qi, M., Zhao, Y., Jia, W., Wang, R.: Audio matters in video super-resolution by implicit semantic guidance · 2022
Closest in time.
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3202–3211 (2022)
Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin transformer · 2022
Closest in time.