Fetching the paper…
Reading the bibliography…
The interactions between human and objects are important for recognizing object-centric actions.
H. W. Kuhn, “The hungarian method for the assignment problem,” Naval research logistics quarterly , vol. 2, no. 1-2, pp. 83–97, 1955
1955
Earlier work this paper cites.
C. Zhang, A. Gupta, and A. Zisserman, “Is an object-centric video representation beneficial for transfer?” in Proceedings of the Asian Conference on Computer Vision , 2022, pp. 1976–1994
1994
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in European conference on computer vision . Springer, Cham, 2016, pp. 20–36
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
X. Zhu, Y. Wang, J. Dai, L. Yuan, and Y. Wei, “Flow-guided feature aggregation for video object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 408–417
2017
Earlier work this paper cites.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The” something something” video database for learning and evaluating visual common sense,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5842–5850
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
D. Li, T. Yao, L.-Y. Duan, T. Mei, and Y. Rui, “Unified spatio-temporal attention networks for action recognition in videos,” IEEE Transactions on Multimedia , vol. 21, no. 2, pp. 416–428, 2018
2018
Earlier work this paper cites.
X. Wang and A. Gupta, “Videos as space-time region graphs,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 399–417
2018
Earlier work this paper cites.
F. Baradel, N. Neverova, C. Wolf, J. Mille, and G. Mori, “Object level visual reasoning in videos,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 105–121
2018
Earlier work this paper cites.
X. Zhu, J. Dai, L. Yuan, and Y. Wei, “Towards high performance video object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7210–7218
2018
Earlier work this paper cites.
S. Wang, Y. Zhou, J. Yan, and Z. Deng, “Fully motion-aware network for video object detection,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 542–557
2018
Earlier work this paper cites.
S. Qi, W. Wang, B. Jia, J. Shen, and S.-C. Zhu, “Learning human-object interactions by graph parsing neural networks,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 401–417
2018
Earlier work this paper cites.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 6202–6211
2019
Cited alongside, same era.
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 7083–7093
2019
Cited alongside, same era.
J. Choi, C. Gao, J. C. Messou, and J.-B. Huang, “Why can’t i dance in the mall? learning to mitigate scene bias in action recognition,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
B. Xu, J. Li, Y. Wong, Q. Zhao, and M. S. Kankanhalli, “Interact as you intend: Intention-driven human-object interaction detection,” IEEE Transactions on Multimedia , vol. 22, no. 6, pp. 1423–1432, 2019
2019
Cited alongside, same era.
T. S. Kim, J. Jones, and G. D. Hager, “Motion guided attention fusion to recognize interactions from videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 076–13 086
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Ben-Shabat, X. Yu, F. Saleh, D. Campbell, C. Rodriguez-Opazo, H. Li, and S. Gould, “The ikea asm dataset: Understanding people assembling furniture through actions, objects and pose,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021, pp. 847–859
2021
Later among the works it cites.
M.-J. Chiou, C.-Y. Liao, L.-W. Wang, R. Zimmermann, and J. Feng, “St-hoi: A spatial-temporal baseline for human-object interaction detection in videos,” in Proceedings of the 2021 Workshop on Intelligent Cross-Data Analysis and Retrieval , pp. 9–17
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Jiang, P. Gao, C. Guo, Q. Zhang, S. Xiang, and C. Pan, “Video object detection with locally-weighted deformable neighbors,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 8529–8536
2019
Cited alongside, same era.
H. Deng, Y. Hua, T. Song, Z. Zhang, Z. Xue, R. Ma, N. Robertson, and H. Guan, “Object guided external memory network for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 6678–6687
2019
Cited alongside, same era.
C. Guo, B. Fan, J. Gu, Q. Zhang, S. Xiang, V. Prinet, and C. Pan, “Progressive sparse local attention for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 3909–3918
2019
Cited alongside, same era.
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 658–666
2019
Cited alongside, same era.
J. Materzynska, T. Xiao, R. Herzig, H. Xu, X. Wang, and T. Darrell, “Something-else: Compositional action recognition with spatial-temporal interaction networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 1049–1059
2020
Cited alongside, same era.
B. Xu, J. Li, Y. Wong, Q. Zhao, and M. S. Kankanhalli, “Interact as you intend: Intention-driven human-object interaction detection,” IEEE Transactions on Multimedia , vol. 22, no. 6, pp. 1423–1432, 2020
2020
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16 . Springer, 2020, pp. 213–229
2020
Cited alongside, same era.
C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 203–213
2020
Cited alongside, same era.
2021
Later among the works it cites.
J. Ji, R. Desai, and J. C. Niebles, “Detecting human-object relationships in videos,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 8086–8096
2021
Later among the works it cites.
M. Tamura, H. Ohashi, and T. Yoshinaga, “Qpic: Query-based pairwise human-object interaction detection with image-wide contextual information,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 10 410–10 419
2021
Later among the works it cites.
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “Mvitv2: Improved multiscale vision transformers for classification and detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4804–4814
2022
Later among the works it cites.
R. Herzig, E. Ben-Avraham, K. Mangalam, A. Bar, G. Chechik, A. Rohrbach, T. Darrell, and A. Globerson, “Object-region video transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3148–3159
2022
Later among the works it cites.
E. Ben Avraham, R. Herzig, K. Mangalam, A. Bar, A. Rohrbach, L. Karlinsky, T. Darrell, and A. Globerson, “Bringing image scene structure to video via frame-clip consistency of object tokens,” Advances in Neural Information Processing Systems , vol. 35, pp. 26 839–26 855, 2022
2022
Later among the works it cites.
Y. Ou, L. Mi, and Z. Chen, “Object-relation reasoning graph for action recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 20 133–20 142
2022
Later among the works it cites.
L. Han and Z. Yin, “Global memory and local continuity for video object detection,” Trans. Multi. , vol. 25, p. 3681–3693, jan 2023. [Online]. Available: https://doi.org/10.1109/TMM.2022.3164253
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Wang, X. Zhang, T. Yang, and J. Sun, “Anchor detr: Query design for transformer-based detector,” in Proceedings of the AAAI conference on artificial intelligence , vol. 36, no. 3, 2022, pp. 2567–2575
2022
Later among the works it cites.
X. Liu, Y.-L. Li, X. Wu, Y.-W. Tai, C. Lu, and C.-K. Tang, “Interactiveness field in human-object interactions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 20 113–20 122
2022
Later among the works it cites.
R. Yan, L. Xie, X. Shu, L. Zhang, and J. Tang, “Progressive instance-aware feature learning for compositional action recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Later among the works it cites.
S. Fang, Z. Lin, K. Yan, J. Li, X. Lin, and R. Ji, “Hodn: Disentangling human-object feature for hoi detection,” Trans. Multi. , vol. 26, p. 3125–3136, aug 2023. [Online]. Available: https://doi.org/10.1109/TMM.2023.3307896
2023
Later among the works it cites.
Q. Qi, Y. Yan, and H. Wang, “Class-aware dual-supervised aggregation network for video object detection,” IEEE Transactions on Multimedia , vol. 26, p. 2109–2123, jul 2023. [Online]. Available: https://doi.org/10.1109/TMM.2023.3292615
2023
Later among the works it cites.
J. Liu, X. Wang, C. Wang, Y. Gao, and M. Liu, “Temporal decoupling graph convolutional network for skeleton-based gesture recognition,” IEEE Transactions on Multimedia , vol. 26, pp. 811–823, 2024
2024
Closest in time.