Fetching the paper…
Reading the bibliography…
Humans have the natural ability to recognize actions even if the objects involved in the action or the background are changed.
Towards real-time object detection with region proposal networks
RCNN Faster · 2015
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Object level visual reasoning in videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Earlier work this paper cites.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2018
Earlier work this paper cites.
ican: Instance-centric attention network for human-object interaction detection
Chen Gao, Yuliang Zou, and Jia-Bin Huang · 2018
Earlier work this paper cites.
Attend and interact: Higher-order object interactions for video understanding
Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, and Hans Peter Graf · 2018
Earlier work this paper cites.
Actor-centric relation network
Chen Sun, Abhinav Shrivastava, Carl Vondrick, Kevin Murphy, Rahul Sukthankar, and Cordelia Schmid · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Videos as space-time region graphs
Xiaolong Wang and Abhinav Gupta · 2018
Earlier work this paper cites.
Graph-based global reasoning networks
Yunpeng Chen, Marcus Rohrbach, Zhicheng Yan, Yan Shuicheng, Jiashi Feng, and Yannis Kalantidis · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Earlier work this paper cites.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 2019
Earlier work this paper cites.
Spatio-temporal action graph networks
Roei Herzig, Elad Levi, Huijuan Xu, Hang Gao, Eli Brosh, Xiaolong Wang, Amir Globerson, and Trevor Darrell · 2019
Earlier work this paper cites.
Transferable interactiveness knowledge for human-object interaction detection
Yong-Lu Li, Siyuan Zhou, Xijie Huang, Liang Xu, Ze Ma, Hao-Shu Fang, Yanfeng Wang, and Cewu Lu · 2019
Earlier work this paper cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Cited alongside, same era.
A structured model for action detection
Yubo Zhang, Pavel Tokmakov, Martial Hebert, and Cordelia Schmid · 2019
Cited alongside, same era.
Drg: Dual relation graph for human-object interaction detection
Chen Gao, Jiarui Xu, Yuliang Zou, and Jia-Bin Huang · 2020
Cited alongside, same era.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Cited alongside, same era.
Uniondet: Union-level detector towards real-time human-object interaction detection
Bumsoo Kim, Taeho Choi, Jaewoo Kang, and Hyunwoo J Kim · 2020
Cited alongside, same era.
Ppdm: Parallel point detection and matching for real-time human-object interaction detection
Detecting human-object relationships in videos
Jingwei Ji, Rishi Desai, and Juan Carlos Niebles · 2021
Later among the works it cites.
Hotr: End-to-end human-object interaction detection with transformers
Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J Kim · 2021
Later among the works it cites.
Motion guided attention fusion to recognize interactions from videos
Tae Soo Kim, Jonathan Jones, and Gregory D Hager · 2021
Later among the works it cites.
Weakly supervised human-object interaction detection in video via contrastive spatiotemporal regions
Shuang Li, Yilun Du, Antonio Torralba, Josef Sivic, and Bryan Russell · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and João F Henriques · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yue Liao, Si Liu, Fei Wang, Yanjie Chen, Chen Qian, and Jiashi Feng · 2020
Cited alongside, same era.
Forecasting human-object interaction: joint prediction of motor attention and actions in first person video
Miao Liu, Siyu Tang, Yin Li, and James M Rehg · 2020
Cited alongside, same era.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell · 2020
Cited alongside, same era.
Knowing what, where and when to look: Efficient video action modeling with attention
Juan-Manuel Perez-Rua, Brais Martinez, Xiatian Zhu, Antoine Toisoul, Victor Escorcia, and Tao Xiang · 2020
Cited alongside, same era.
Spatio-temporal action detection with multi-object interaction
Huijuan Xu, Lizhi Yang, Stan Sclaroff, Kate Saenko, and Trevor Darrell · 2020
Cited alongside, same era.
Interactive fusion of multi-level features for compositional activity recognition
Rui Yan, Lingxi Xie, Xiangbo Shu, and Jinhui Tang · 2020
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Gorjan Radevski, Marie-Francine Moens, and Tinne Tuytelaars · 2021
Later among the works it cites.
Mining the benefits of two-stage and one-stage hoi detection
Aixi Zhang, Yue Liao, Si Liu, Miao Lu, Yongliang Wang, Chen Gao, and Xiaobo Li · 2021
Later among the works it cites.
Object-region video transformers
Roei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar, Gal Chechik, Anna Rohrbach, Trevor Darrell, and Amir Globerson · 2022
Later among the works it cites.
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xiaolong Wang · 2022
Later among the works it cites.
Object-relation reasoning graph for action recognition
Yangjun Ou, Li Mi, and Zhenzhong Chen · 2022
Later among the works it cites.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Later among the works it cites.
Is an object-centric video representation beneficial for transfer?
Chuhan Zhang, Ankush Gupta, and Andrew Zisserman · 2022
Later among the works it cites.
Look, remember and reason: Grounded reasoning in videos with language models
Apratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee, Pulkit Madan, and Roland Memisevic · 2023
Closest in time.
Semantic-disentangled transformer with noun-verb embedding for compositional action recognition
Peng Huang, Rui Yan, Xiangbo Shu, Zhewei Tu, Guangzhao Dai, and Jinhui Tang · 2023
Closest in time.
Progressive instance-aware feature learning for compositional action recognition
Rui Yan, Lingxi Xie, Xiangbo Shu, Liyan Zhang, and Jinhui Tang · 2023
Closest in time.
Interaction region visual transformer for egocentric action anticipation
Debaditya Roy, Ramanathan Rajendiran, and Basura Fernando · 2024
Closest in time.