Fetching the paper…
Reading the bibliography…
First-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention.
Miller, G.: Wordnet: a lexical database for english. In: CACM (1995)
1995
Earlier work this paper cites.
Banerjee, S., Pedersen, T.: An adapted lesk algorithm for word sense disambiguation using wordnet. In: CICLing (2002)
2002
Earlier work this paper cites.
Zach, C., Pock, T., Bischof, H.: A duality based approach for realtime TV-L1 optical flow. In: Pattern Recognition (2007)
2007
Earlier work this paper cites.
De La Torre, F., Hodgins, J., Bargteil, A., Martin, X., Macey, J., Collado, A., Beltran, P.: Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) database. In: Robotics Institute (2008)
2008
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR (2009)
2009
Earlier work this paper cites.
Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: The PASCAL Visual Object Classes (VOC) Challenge. In: IJCV (2010)
2010
Earlier work this paper cites.
Fathi, A., Hodgins, J., Rehg, J.: Social interactions: A first-person perspective. In: CVPR (2012)
2012
Earlier work this paper cites.
Fathi, A., Li, Y., Rehg, J.: Learning to recognize daily actions using gaze. In: ECCV (2012)
2012
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
2012
Earlier work this paper cites.
Lee, Y., Ghosh, J., Grauman, K.: Discovering important people and objects for egocentric video summarization. In: CVPR (2012)
2012
Earlier work this paper cites.
Pirsiavash, H., Ramanan, D.: Detecting activities of daily living in first-person camera views. In: CVPR (2012)
2012
Earlier work this paper cites.
Rohrbach, M., Amin, S., Andriluka, M., Schiele, B.: A Database for Fine Grained Activity Detection of Cooking Activities. In: CVPR (2012)
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
Ryoo, M.S., Matthies, L.: First-person activity recognition: What are they doing to me? In: CVPR (2013)
2013
Earlier work this paper cites.
Stein, S., McKenna, S.: Combining Embedded Accelerometers with Computer Vision for Recognizing Food Preparation Activities. In: UbiComp (2013)
2013
Earlier work this paper cites.
Damen, D., Leelasawassuk, T., Haines, O., Calway, A., Mayol-Cuevas, W.: You-do, I-learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video. In: BMVC (2014)
2014
Earlier work this paper cites.
Kuehne, H., Arslan, A., Serre, T.: The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities. In: CVPR (2014)
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014)
2014
Cited alongside, same era.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Advances in neural information processing systems. pp. 568–576 (2014)
2014
Cited alongside, same era.
Alletto, S., Serra, G., Calderara, S., Cucchiara, R.: Understanding social relationships in egocentric vision. In: Pattern Recognition (2015)
2015
Cited alongside, same era.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: VQA: Visual Question Answering. In: ICCV (2015)
2015
Cited alongside, same era.
Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: ICML (2015)
2017
Later among the works it cites.
Furnari, A., Battiato, S., Grauman, K., Farinella, G.M.: Next-active-object prediction from egocentric videos. In: JVCIR (2017)
2017
Later among the works it cites.
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fründ, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.: The ”something something” video database for learning and evaluating visual common sense. In: ICCV (2017)
2017
Later among the works it cites.
Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., Fischer, I., Wojna, Z., Song, Y., Guadarrama, S., et al.: Speed/accuracy trade-offs for modern convolutional object detectors. In: CVPR (2017)
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Karpathy, A., Fei-Fei, L.: Deep Visual-Semantic Alignments for Generating Image Descriptions. In: CVPR (2015)
2015
Cited alongside, same era.
Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. In: NIPS (2015)
2015
Cited alongside, same era.
Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: A Dataset for Movie Description. In: CVPR (2015)
2015
Cited alongside, same era.
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In: CVPR (2015)
2015
Cited alongside, same era.
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., Vijayanarasimhan, S.: YouTube-8M: A Large-Scale Video Classification Benchmark. In: CoRR (2016)
2016
Cited alongside, same era.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
2016
Cited alongside, same era.
Park, H.S., Hwang, J.J., Niu, Y., Shi, J.: Egocentric future localization. In: CVPR (2016)
2016
Cited alongside, same era.
Kalogeiton, V., Weinzaepfel, P., Ferrari, V., Schmid, C.: Joint learning of object and action detectors. In: ICCV (2017)
2017
Later among the works it cites.
Moltisanti, D., Wray, M., Mayol-Cuevas, W., Damen, D.: Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video. In: ICCV (2017)
2017
Later among the works it cites.
Nair, A., Chen, D., Agrawal, P., Isola, P., Abbeel, P., Malik, J., Levine, S.: Combining self-supervised learning and imitation for vision-based rope manipulation. In: ICRA (2017)
2017
Later among the works it cites.
Yuanjun, X.: PyTorch Temporal Segment Network. https://github.com/yjxiong/tsn-pytorch (2017)
2017
Later among the works it cites.
2017
Later among the works it cites.
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Scene parsing through ade20k dataset. In: CVPR (2017)
2017
Later among the works it cites.
2017
Later among the works it cites.
Doughty, H., Damen, D., Mayol-Cuevas, W.: Who’s better? who’s best? pairwise deep ranking for skill determination. In: CVPR (2018)
2018
Closest in time.
Georgia Tech: Extended GTEA Gaze+. http://webshare.ipat.gatech.edu/coc-rim-wall-lab/web/yli440/egtea_gp (2018)
2018
Closest in time.
Heidarivincheh, F., Mirmehdi, M., Damen, D.: Action completion: A temporal model for moment detection. In: BMVC (2018)
2018
Closest in time.
Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., Alahari, K.: Charades-ego: A large-scale dataset of paired third and first person videos. In: ArXiv (2018)
2018
Closest in time.
Yeung, S., Russakovsky, O., Jin, N., Andriluka, M., Mori, G., Fei-Fei, L.: Every moment counts: Dense detailed labeling of actions in complex videos. IJCV (2018)
2018
Closest in time.
Zhang, T., McCarthy, Z., Jow, O., Lee, D., Goldberg, K., Abbeel, P.: Deep imitation learning for complex manipulation tasks from virtual reality teleoperation. In: ICRA (2018)
2018
Closest in time.