Fetching the paper…
Reading the bibliography…
Hands are the central means by which humans manipulate their world and being able to reliably extract hand state information from Internet videos of humans engaged in their hands has the potential to pave the way to systems that can learn from petabytes of video data.
The ecological approach to visual perception
J. Gibson · 1979
Earlier work this paper cites.
The PASCAL Visual Object Classes (VOC) challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
An object-dependent hand pose prior from sparse training data
Henning Hamer, Juergen Gall, Thibaut Weise, and Luc Van Gool · 2010
Earlier work this paper cites.
Hand detection using multiple proposals
A. Mittal, A. Zisserman, and P. H. S. Torr · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. A. Efros · 2011
Earlier work this paper cites.
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions
Sven Bambach, Stefan Lee, David Crandall, and Chen Yu · 2015
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
Saurabh Gupta and Jitendra Malik · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Understanding everyday hands in action from rgb-d images
G. Rogez, J. Supancic, and D. Ramanan · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Real-time joint tracking of a hand manipulating an object from rgb-d input
Srinath Sridhar, Franziska Mueller, Michael Zollhöfer, Dan Casas, Antti Oulasvirta, and Christian Theobalt · 2016
Cited alongside, same era.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Cited alongside, same era.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Cited alongside, same era.
Embodied hands: Modeling and capturing hands and bodies together
J. Romero, D. Tzionas, and M. J. Black · 2017
Cited alongside, same era.
Detecting and recognizing human-object interactions
Georgia Gkioxari, Ross Girshick, Piotr Dollar, and Kaiming He · 2018
Later among the works it cites.
AVA: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, Sudheendra Vijayanarasimhan, Caroline Pantofaru, David A. Ross, George Toderici, Yeqing Li, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik · 2018
Later among the works it cites.
Cross-modal deep variational hand pose estimation
A. Spurr, J. Song, S. Park, and O. Hilliges · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2018
Later among the works it cites.
Contactdb: Analyzing and predicting grasp contact via thermal imaging
Samarth Brahmbhatt, Cusuh Ham, Charlie Kemp, and James Hays · 2019
Later among the works it cites.
Learning joint reconstruction of hands and manipulated objects
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Zimmerman and T. Brox · 2017
Cited alongside, same era.
Tactile logging for understanding plausible tool use based on human demonstration
Shuichi Akizuki and Yoshimitsu Aoki · 2018
Cited alongside, same era.
Confidence from invariance to image transformations
Yuval Bahat and Gregory Shakhnarovich · 2018
Cited alongside, same era.
Object level visual reasoning in videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Cited alongside, same era.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2018
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Cited alongside, same era.
Demo2vec: Reasoning object affordances from online videos
Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese, and Joseph J. Lim · 2018
Cited alongside, same era.
Yana Hasson, Gül Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfruend, Carl Vondrick, et al · 2019
Later among the works it cites.
Grounded human-object interaction hotspots from video
Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman · 2019
Later among the works it cites.
Contextual attention for hand detection in the wild
Supreeth Narasimhaswamy, Zhengwei Wei, Yang Wang, Justin Zhang, and Minh Hoai · 2019
Later among the works it cites.
H+o: Unified egocentric recognition of 3d hand-object poses and interactions
Bugra Tekin, Federica Bobo, and Marc Pollefeys · 2019
Later among the works it cites.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox · 2019
Later among the works it cites.