Fetching the paper…
Reading the bibliography…
We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera.
B. Bridgeman, D. Hendry, and L. Stark, “Failure to detect displacement of the visual world during saccadic eye movements,” Vision research , vol. 15, no. 6, pp. 719–722, 1975
1975
Earlier work this paper cites.
L. Itti and C. Koch, “Computational modelling of visual attention,” Nature reviews neuroscience , vol. 2, no. 3, p. 194, 2001
2001
Earlier work this paper cites.
J. M. Henderson, “Human gaze control during real-world scene perception,” Trends in cognitive sciences , vol. 7, no. 11, pp. 498–504, 2003
2003
Earlier work this paper cites.
P. Wittenburg, H. Brugman, A. Russel, A. Klassmann, and H. Sloetjes, “ELAN: a professional framework for multimodality research,” in Proceedings of LREC , vol. 2006, 2006, p. 5th
2006
Earlier work this paper cites.
P. Turaga, R. Chellappa, V. S. Subrahmanian, and O. Udrea, “Machine recognition of human activities: A survey,” Circuits and Systems for Video Technology, IEEE Transactions on , vol. 18, no. 11, pp. 1473–1488, Nov 2008
2008
Earlier work this paper cites.
F. D. la Torre Frade, J. K. Hodgins, A. W. Bargteil, X. M. Artal, J. C. Macey, A. C. I. Castells, and J. Beltran, “Guide to the carnegie mellon university multimodal activity (CMU-MMAC) database,” Carnegie Mellon University, Pittsburgh, PA, Tech. Rep. CMU-RI-TR-08-22, April 2008
2008
Earlier work this paper cites.
E. H. Spriggs, F. De La Torre, and M. Hebert, “Temporal segmentation and activity classification from first-person sensing,” in CVPR Workshops , 2009
2009
Earlier work this paper cites.
D. W. Hansen and Q. Ji, “In the eye of the beholder: A survey of models for eyes and gaze,” TPAMI , vol. 32, no. 3, pp. 478–500, 2010
2010
Earlier work this paper cites.
F. Perronnin, J. Sánchez, and T. Mensink, “Improving the fisher kernel for large-scale image classification,” in ECCV , 2010
2010
Earlier work this paper cites.
K. M. Kitani, T. Okabe, Y. Sato, and A. Sugimoto, “Fast unsupervised ego-action learning for first-person sports videos,” in CVPR , 2011
2011
Earlier work this paper cites.
J. Aggarwal and M. Ryoo, “Human activity analysis: A review,” ACM Comput. Surv. , vol. 43, no. 3, pp. 16:1–16:43, Apr 2011
2011
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “HMDB: a large video database for human motion recognition,” in ICCV , 2011
2011
Earlier work this paper cites.
A. Fathi, A. Farhadi, and J. M. Rehg, “Understanding egocentric activities,” in ICCV , 2011
2011
Earlier work this paper cites.
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu, “Action recognition by dense trajectories,” in CVPR , 2011
2011
Earlier work this paper cites.
T. Kanade and M. Hebert, “First-person vision,” Proceedings of the IEEE , vol. 100, no. 8, pp. 2442 –2453, 2012
2012
Earlier work this paper cites.
H. Pirsiavash and D. Ramanan, “Detecting activities of daily living in first-person camera views,” in CVPR , 2012
2012
Earlier work this paper cites.
H. S. Park, E. Jain, and Y. Sheikh, “3D social saliency from head-mounted cameras.” in NeurIPS , 2012
2012
Earlier work this paper cites.
A. Fathi, Y. Li, and J. M. Rehg, “Learning to recognize daily actions using gaze,” in ECCV , 2012
2012
Earlier work this paper cites.
S. Mathe and C. Sminchisescu, “Dynamic eye movement datasets and learnt saliency models for visual action recognition,” in ECCV , 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele, “A database for fine grained activity detection of cooking activities,” in CVPR , 2012
2012
Earlier work this paper cites.
A. Fathi, J. K. Hodgins, and J. M. Rehg, “Social interactions: A first-person perspective,” in CVPR , 2012
2012
Earlier work this paper cites.
M. S. Ryoo and L. Matthies, “First-person activity recognition: What are they doing to me?” in CVPR , 2013
2013
Earlier work this paper cites.
Y. Li, A. Fathi, and J. M. Rehg, “Learning to predict gaze in egocentric video,” in ICCV , 2013
2013
Earlier work this paper cites.
N. Shapovalova, M. Raptis, L. Sigal, and G. Mori, “Action is in the eye of the beholder: Eye-gaze driven model for spatio-temporal action localization,” in NeurIPS , 2013
2013
Earlier work this paper cites.
C. Li and K. M. Kitani, “Model recommendation with virtual probes for egocentric hand detection,” in ICCV , 2013
2013
Earlier work this paper cites.
A. Borji and L. Itti, “State-of-the-art in visual attention modeling,” TPAMI , vol. 35, no. 1, pp. 185–207, 2013
2013
Earlier work this paper cites.
S. Bell, P. Upchurch, N. Snavely, and K. Bala, “OpenSurfaces: A richly annotated catalog of surface appearance,” ACM Transactions on Graphics (TOG) , vol. 32, no. 4, p. 111, 2013
2013
Earlier work this paper cites.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in CVPR , 2014
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NeurIPS , 2014
2014
Earlier work this paper cites.
W. Chen, C. Xiong, R. Xu, and J. J. Corso, “Actionness ranking with lattice conditional ordinal random fields,” in CVPR , 2014
2014
Earlier work this paper cites.
J. Hernandez, Y. Li, J. M. Rehg, and R. W. Picard, “Bioglass: Physiological parameter estimation using a head-mounted wearable device,” in Wireless Mobile Communication and Healthcare (Mobihealth), 2014 EAI 4th International Conference on . IEEE, 2014
2014
Earlier work this paper cites.
Y. Poleg, C. Arora, and S. Peleg, “Head motion signatures from egocentric videos,” in ACCV , 2014
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR , 2014
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in ICLR , 2014
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
Y. Li, Z. Ye, and J. M. Rehg, “Delving into egocentric actions,” in CVPR , 2015
2015
Earlier work this paper cites.
H. S. Park and J. Shi, “Social saliency prediction,” in CVPR , 2015
2015
Earlier work this paper cites.
G. Rogez, J. S. Supancic, III, and D. Ramanan, “First-person pose recognition using egocentric workspaces,” in CVPR , 2015
2015
Earlier work this paper cites.
J. Xu, L. Mukherjee, Y. Li, J. Warner, J. M. Rehg, and V. Singh, “Gaze-enabled egocentric video summarization via constrained submodular maximization,” in CVPR , 2015
2015
Cited alongside, same era.
Y. J. Lee and K. Grauman, “Predicting important objects for egocentric video summarization,” IJCV , vol. 114, no. 1, pp. 38–55, 2015
2015
Cited alongside, same era.
Y. Poleg, T. Halperin, C. Arora, and S. Peleg, “Egosampling: Fast-forward and stereo for egocentric videos,” in CVPR , 2015
2015
Cited alongside, same era.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015
2015
Cited alongside, same era.
L. Wang, Y. Qiao, and X. Tang, “Action recognition with trajectory-pooled deep-convolutional descriptors,” in CVPR , 2015
M. Zhang, K. Teck Ma, J. Hwee Lim, Q. Zhao, and J. Feng, “Deep future gaze: Gaze anticipation on egocentric videos using adversarial networks,” in CVPR , 2017
2017
Later among the works it cites.
G. A. Sigurdsson, O. Russakovsky, and A. Gupta, “What actions are needed for understanding human actions in videos?” in ICCV , 2017
2017
Later among the works it cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in ICLR , 2017
2017
Later among the works it cites.
C. J. Maddison, A. Mnih, and Y. W. Teh, “The concrete distribution: A continuous relaxation of discrete random variables,” in ICLR , 2017
2017
Later among the works it cites.
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox, “Flownet 2.0: Evolution of optical flow estimation with deep networks,” in ICCV , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
G. Cheron, I. Laptev, and C. Schmid, “P-cnn: Pose-based cnn features for action recognition,” in ICCV , December 2015
2015
Cited alongside, same era.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in ICCV , 2015
2015
Cited alongside, same era.
D.-A. Huang, M. Ma, W.-C. Ma, and K. M. Kitani, “How do we use our hands? discovering a diverse set of common grasps,” in CVPR , 2015
2015
Cited alongside, same era.
G. Rogez, J. S. Supancic, III, and D. Ramanan, “Understanding everyday hands in action from rgb-d images,” in ICCV , 2015
2015
Cited alongside, same era.
R. Yonetani, K. M. Kitani, and Y. Sato, “Ego-surfing first-person videos,” in CVPR , 2015
2015
Cited alongside, same era.
A. Betancourt, P. Morerio, C. S. Regazzoni, and M. Rauterberg, “The evolution of first person vision methods: A survey,” Circuits and Systems for Video Technology, IEEE Transactions on , vol. 25, no. 5, pp. 744–760, 2015
2015
Cited alongside, same era.
M. S. Ryoo, B. Rothrock, and L. Matthies, “Pooled motion features for first-person videos,” in CVPR , 2015
2015
Cited alongside, same era.
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray, “Scaling egocentric vision: The EPIC-KITCHENS dataset,” in ECCV , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Li, M. Liu, and J. M. Rehg, “In the eye of beholder: Joint learning of gaze and actions in first person video,” in ECCV , 2018
2018
Later among the works it cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in CVPR , 2018
2018
Later among the works it cites.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in ECCV , 2018
2018
Later among the works it cites.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in CVPR , 2018
2018
Later among the works it cites.
Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng, “Multi-fiber networks for video recognition,” in ECCV , 2018
2018
Later among the works it cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in ECCV , 2018
2018
Later among the works it cites.
D. Wei, J. J. Lim, A. Zisserman, and W. T. Freeman, “Learning and using the arrow of time,” in CVPR , 2018
2018
Later among the works it cites.
X. Wang, R. B. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018
2018
Later among the works it cites.
Y. Huang, M. Cai, Z. Li, and Y. Sato, “Predicting gaze in egocentric video by learning task-dependent attention transition,” in ECCV , 2018
2018
Later among the works it cites.
Y. Shen, B. Ni, Z. Li, and N. Zhuang, “Egocentric activity prediction via event modulated attention,” in ECCV , 2018
2018
Later among the works it cites.
S. Sudhakaran and O. Lanz, “Attention is all we need: Nailing down object-centric attention for egocentric activity recognition,” in BMVC , 2018
2018
Later among the works it cites.
G. Ghiasi, T.-Y. Lin, and Q. V. Le, “Dropblock: A regularization method for convolutional networks,” in NeurIPS , 2018
2018
Later among the works it cites.
S. Sudhakaran, S. Escalera, and O. Lanz, “LSTA: Long short-term attention for egocentric action recognition,” in CVPR , 2019
2019
Later among the works it cites.
A. Furnari and G. M. Farinella, “What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention,” in ICCV , 2019
2019
Later among the works it cites.
C. Luo and A. L. Yuille, “Grouped spatial-temporal aggregation for efficient action recognition,” in ICCV , 2019
2019
Later among the works it cites.
D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in ICCV , 2019
2019
Later among the works it cites.
A. Piergiovanni, A. Angelova, A. Toshev, and M. S. Ryoo, “Evolving space-time neural architectures for videos,” in ICCV , 2019
2019
Later among the works it cites.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in ICCV , 2019
2019
Later among the works it cites.
N. Crasto, P. Weinzaepfel, K. Alahari, and C. Schmid, “MARS: Motion-Augmented RGB Stream for Action Recognition,” in CVPR , 2019
2019
Later among the works it cites.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in ICCV , 2019
2019
Later among the works it cites.
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick, “Long-term feature banks for detailed video understanding,” in CVPR , 2019
2019
Later among the works it cites.
B. Tekin, F. Bogo, and M. Pollefeys, “H+O: Unified egocentric recognition of 3d hand-object poses and interactions,” in CVPR , 2019
2019
Later among the works it cites.
H. Li, Y. Cai, and W.-S. Zheng, “Deep dual relation modeling for egocentric interaction recognition,” in CVPR , 2019
2019
Later among the works it cites.
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “EPIC-Fusion: Audio-visual temporal binding for egocentric action recognition,” in ICCV , 2019
2019
Later among the works it cites.
M. Wray, D. Larlus, G. Csurka, and D. Damen, “Fine-grained action retrieval through multiple parts-of-speech embeddings,” in ICCV , 2019
2019
Later among the works it cites.
G. Kapidis, R. Poppe, E. van Dam, L. Noldus, and R. Veltkamp, “Multitask learning to improve egocentric action recognition,” in ICCV Workshops , 2019
2019
Later among the works it cites.
D. Ghadiyaram, D. Tran, and D. Mahajan, “Large-scale weakly-supervised pre-training for video action recognition,” in CVPR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Huang, M. Cai, Z. Li, F. Lu, and Y. Sato, “Mutual context network for jointly estimating egocentric gaze and action,” TIP , vol. 29, pp. 7795–7806, 2020
2020
Closest in time.
A. Bandini and J. Zariffa, “Analysis of the hands in egocentric vision: A survey,” TPAMI , 2020
2020
Closest in time.
M. Liu, S. Tang, Y. Li, and J. Rehg, “Forecasting human object interaction: Joint prediction of motor attention and egocentric activity,” in ECCV , 2020
2020
Closest in time.
M. Liu, X. Chen, Y. Zhang, Y. Li, and J. M. Rehg, “Attention distillation for learning video representations,” in BMVC , 2020
2020
Closest in time.