Fetching the paper…
Reading the bibliography…
We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects.
The unscented kalman filter for nonlinear estimation
Eric A Wan and Rudolph Van Der Merwe · 2000
Earlier work this paper cites.
The recognition of human movement using temporal templates
Aaron F. Bobick and James W. Davis · 2001
Earlier work this paper cites.
” grabcut” interactive foreground extraction using iterated graph cuts
Carsten Rother, Vladimir Kolmogorov, and Andrew Blake · 2004
Earlier work this paper cites.
Behavior recognition via sparse spatio-temporal features
Piotr Dollár, Vincent Rabaud, Garrison Cottrell, and Serge Belongie · 2005
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
A biologically inspired system for action recognition
Hueihan Jhuang, Thomas Serre, Lior Wolf, and Tomaso Poggio · 2007
Earlier work this paper cites.
Learning motion categories using both semantic and structural information
Shu-Fai Wong, Tae-Kyun Kim, and Roberto Cipolla · 2007
Earlier work this paper cites.
Unsupervised learning of human action categories using spatial-temporal words
Juan Carlos Niebles, Hongcheng Wang, and Li Fei-Fei · 2008
Earlier work this paper cites.
Epnp: An accurate o (n) solution to the pnp problem
Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua · 2009
Earlier work this paper cites.
EPnP: An Accurate O(n) Solution to the PnP Problem
V. Lepetit, F. Moreno-Noguer, and P. Fua · 2009
Earlier work this paper cites.
High level activity recognition using low resolution wearable vision
Sudeep Sundaram and Walterio W Mayol Cuevas · 2009
Earlier work this paper cites.
Understanding egocentric activities
Alireza Fathi, Ali Farhadi, and James M Rehg · 2011
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Alireza Fathi, Xiaofeng Ren, and James M Rehg · 2011
Earlier work this paper cites.
Fast unsupervised ego-action learning for first-person sports videos
Kris M Kitani, Takahiro Okabe, Yoichi Sato, and Akihiro Sugimoto · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu · 2011
Earlier work this paper cites.
Motion capture of hands in action using discriminative salient points
Luca Ballan, Aparna Taneja, Jürgen Gall, Luc Van Gool, and Marc Pollefeys · 2012
Earlier work this paper cites.
Learning to recognize daily actions using gaze
Alireza Fathi, Yin Li, and James M Rehg · 2012
Earlier work this paper cites.
Spatio-temporal human-object interactions for action recognition in videos
Victor Escorcia and Juan Niebles · 2013
Earlier work this paper cites.
Learning to predict gaze in egocentric video
Yin Li, Alireza Fathi, and James M Rehg · 2013
Earlier work this paper cites.
First-person activity recognition: What are they doing to me?
Michael S Ryoo and Larry Matthies · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Capturing hand motion with an rgb-d sensor, fusing a generative model with salient points
Dimitrios Tzionas, Abhilash Srikantha, Pablo Aponte, and Juergen Gall · 2014
Earlier work this paper cites.
Panoptic studio: A massively multiview system for social motion capture
Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Delving into egocentric actions
Yin Li, Zhefan Ye, and James M Rehg · 2015
Earlier work this paper cites.
First person action-object detection with egonet
Gedas Bertasius, Hyun Soo Park, Stella X Yu, and Jianbo Shi · 2016
Earlier work this paper cites.
Uncertainty-driven 6d pose estimation of objects and scenes from a single rgb image
Eric Brachmann, Frank Michel, Alexander Krull, Michael Ying Yang, Stefan Gumhold, et al · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Going deeper into first-person activity recognition
Minghuang Ma, Haoqi Fan, and Kris M Kitani · 2016
Earlier work this paper cites.
Hand-object interaction and precise localization in transitive action recognition
Amir Rosenfeld and Shimon Ullman · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
First person action recognition using deep learned descriptors
Suriya Singh, Chetan Arora, and CV Jawahar · 2016
Earlier work this paper cites.
Real-time joint tracking of a hand manipulating an object from rgb-d input
Srinath Sridhar, Franziska Mueller, Michael Zollhoefer, Dan Casas, Antti Oulasvirta, and Christian Theobalt · 2016
Earlier work this paper cites.
Capturing hands in action using discriminative salient pointsand physics simulation
Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha, Pablo Aponte, Marc Pollefeys, and Juergen Gall · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Spatial attention deep net with partial pso for hierarchical hybrid hand pose estimation
Qi Ye, Shanxin Yuan, and Tae-Kyun Kim · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Later among the works it cites.
ContactDB: Analyzing and predicting grasp contact via thermal imaging
Samarth Brahmbhatt, Cusuh Ham, Charles C. Kemp, and James Hays · 2019
Later among the works it cites.
Openpose: realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2019
Later among the works it cites.
The VIA annotation software for images, audio and video
Abhishek Dutta and Andrew Zisserman · 2019
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
Panoptic studio: A massively multiview system for social interaction capture
Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Scott Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Real-time hand tracking under occlusion from an egocentric rgb-d sensor
Franziska Mueller, Dushyant Mehta, Oleksandr Sotnychenko, Srinath Sridhar, Dan Casas, and Christian Theobalt · 2017
Cited alongside, same era.
Deepprior++: Improving fast and accurate 3d hand pose estimation
Markus Oberweger and Vincent Lepetit · 2017
Cited alongside, same era.
YOLO9000: Better, Faster, Stronger
J. Redmon and A. Farhadi · 2017
Cited alongside, same era.
Embodied hands: Modeling and capturing hands and bodies together
Javier Romero, Dimitrios Tzionas, and Michael J Black · 2017
Cited alongside, same era.
Later among the works it cites.
No-frills human-object interaction detection: Factorization, layout encodings, and training techniques
Tanmay Gupta, Alexander Schwing, and Derek Hoiem · 2019
Later among the works it cites.
Learning joint reconstruction of hands and manipulated objects
Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid · 2019
Later among the works it cites.
Actional-structural graph convolutional networks for skeleton-based action recognition
Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian · 2019
Later among the works it cites.
Transferable interactiveness knowledge for human-object interaction detection
Yong-Lu Li, Siyuan Zhou, Xijie Huang, Liang Xu, Ze Ma, Hao-Shu Fang, Yanfeng Wang, and Cewu Lu · 2019
Later among the works it cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
Pvnet: Pixel-wise voting network for 6dof pose estimation
Sida Peng, Yuan Liu, Qixing Huang, Xiaowei Zhou, and Hujun Bao · 2019
Later among the works it cites.
Bad slam: Bundle adjusted direct rgb-d slam
Thomas Schops, Torsten Sattler, and Marc Pollefeys · 2019
Later among the works it cites.
Two-stream adaptive graph convolutional networks for skeleton-based action recognition
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu · 2019
Later among the works it cites.
H+ o: Unified egocentric recognition of 3d hand-object poses and interactions
Bugra Tekin, Federica Bogo, and Marc Pollefeys · 2019
Later among the works it cites.
Pose-aware multi-level feature network for human object interaction detection
Bo Wan, Desen Zhou, Yongfei Liu, Rongjie Li, and Xuming He · 2019
Later among the works it cites.
Densefusion: 6d object pose estimation by iterative dense fusion
Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martín-Martín, Cewu Lu, Li Fei-Fei, and Silvio Savarese · 2019
Later among the works it cites.
Deep contextual attention for human-object interaction detection
Tiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao, and Jorma Laaksonen · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Later among the works it cites.
Reasoning about human-object interactions through dual attention networks
Tete Xiao, Quanfu Fan, Dan Gutfreund, Mathew Monfort, Aude Oliva, and Bolei Zhou · 2019
Later among the works it cites.
Relation parsing neural network for human-object interaction detection
Penghao Zhou and Mingmin Chi · 2019
Later among the works it cites.
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox · 2019
Later among the works it cites.
ContactPose: A dataset of grasps with object contact and hand pose
Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg, Charles C. Kemp, and James Hays · 2020
Later among the works it cites.
Skeleton-based action recognition with shift graph convolutional network
Ke Cheng, Yifan Zhang, Xiangyu He, Weihan Chen, Jian Cheng, and Hanqing Lu · 2020
Later among the works it cites.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Later among the works it cites.
Pyslowfast
Haoqi Fan, Yanghao Li, Bo Xiong, Wan-Yen Lo, and Christoph Feichtenhofer · 2020
Later among the works it cites.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Later among the works it cites.
Honnotate: A method for 3d annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit · 2020
Later among the works it cites.
Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction
Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, and Cordelia Schmid · 2020
Later among the works it cites.
Disentangling and unifying graph convolutions for skeleton-based action recognition
Ziyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang, and Wanli Ouyang · 2020
Later among the works it cites.
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee · 2020
Later among the works it cites.
Francesco Ragusa, Antonino Furnari, Salvatore Livatino, and Giovanni Maria Farinella · 2020
Later among the works it cites.
GRAB: A dataset of whole-body human grasping of objects
Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas · 2020
Later among the works it cites.
Azure Kinect SDK
Azure SDK Team · 2020
Later among the works it cites.