Fetching the paper…
Reading the bibliography…
Since its introduction in 2018, EPIC-KITCHENS has attracted attention as the largest egocentric video benchmark, offering a unique viewpoint on people's interaction with objects, their attention, and even intention.
G. Miller, “Wordnet: a lexical database for english,” in
1995
Earlier work this paper cites.
S. Banerjee and T. Pedersen, “An adapted lesk algorithm for word sense disambiguation using wordnet,” in
2002
Earlier work this paper cites.
C. Zach, T. Pock, and H. Bischof, “A duality based approach for realtime TV-L1 optical flow,” in
2007
Earlier work this paper cites.
C. Zach, T. Pock, and H. Bischof, “A duality based approach for realtime TV-L1 optical flow,” in
2007
Earlier work this paper cites.
F. De La Torre, J. Hodgins, A. Bargteil, X. Martin, J. Macey, A. Collado, and P. Beltran, “Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) database,” in
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
Y. Xu and D Damen. “Human Routine Change Detection using Bayesian Modelling”. In
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes (VOC) Challenge,” in
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
A. Fathi, Y. Li, and J. Rehg, “Learning to recognize daily actions using gaze,” in
2012
Earlier work this paper cites.
H. Pirsiavash and D. Ramanan, “Detecting activities of daily living in first-person camera views,” in
2012
Earlier work this paper cites.
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele, “A Database for Fine Grained Activity Detection of Cooking Activities,” in
2012
Earlier work this paper cites.
A. Fathi, J. Hodgins, and J. Rehg, “Social interactions: A first-person perspective,” in
2012
Earlier work this paper cites.
Y. Lee, J. Ghosh, and K. Grauman, “Discovering important people and objects for egocentric video summarization,” in
2012
Earlier work this paper cites.
S. Stein and S. McKenna, “Combining Embedded Accelerometers with Computer Vision for Recognizing Food Preparation Activities,” in
2013
Earlier work this paper cites.
M. S. Ryoo and L. Matthies, “First-person activity recognition: What are they doing to me?” in
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in
2014
Earlier work this paper cites.
D. Damen, T. Leelasawassuk, O. Haines, A. Calway, and W. Mayol-Cuevas, “You-do, I-learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video,” in
2014
Earlier work this paper cites.
H. Kuehne, A. Arslan, and T. Serre, “The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities,” in
2014
Cited alongside, same era.
K. Simonyan, and A. Zisserman. “Two-stream convolutional networks for action recognition in videos,” in
2014
Cited alongside, same era.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in
2014
Cited alongside, same era.
Simonyan, Karen, and Andrew Zisserman. “Two-stream convolutional networks for action recognition in videos,” in
2014
Cited alongside, same era.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Later among the works it cites.
C. Vondrick, H. Pirsiavash, and A. Torralba, “Anticipating visual representations from unlabeled video,” in
2016
Later among the works it cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in
2017
Later among the works it cites.
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fründ, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic, “The ”something something” video database for learning and evaluating visual common sense,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Karpathy and L. Fei-Fei, “Deep Visual-Semantic Alignments for Generating Image Descriptions,” in
2015
Cited alongside, same era.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in
2015
Cited alongside, same era.
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele, “A Dataset for Movie Description,” in
2015
Cited alongside, same era.
S. Alletto, G. Serra, S. Calderara, and R. Cucchiara, “Understanding social relationships in egocentric vision,” in
2015
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in
2015
Cited alongside, same era.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Cited alongside, same era.
2017
Later among the works it cites.
J.-B. Alayrac, J. Sivic, I. Laptev, and S. Lacoste-Julien, “Joint discovery of object states and manipulation actions,” in
2017
Later among the works it cites.
D. Moltisanti, M. Wray, W. Mayol-Cuevas, and D. Damen, “Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video,” in
2017
Later among the works it cites.
V. Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid, “Joint learning of object and action detectors,” in
2017
Later among the works it cites.
A. Furnari, S. Battiato, K. Grauman, and G. M. Farinella, “Next-active-object prediction from egocentric videos,” in
2017
Later among the works it cites.
N. Rhinehart, and K. Kitani. “First-person activity forecasting with online inverse reinforcement learning,” in
2017
Later among the works it cites.
J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama and K. Murphy, “Speed/accuracy trade-offs for modern convolutional object detectors,” in
2017
Later among the works it cites.
X. Yuanjun, “PyTorch Temporal Segment Network,” https://github.com/yjxiong/tsn-pytorch, 2017
2017
Later among the works it cites.
D. F. Fouhey, W.-c. Kuo, A. A. Efros, and J. Malik, “From lifestyle vlogs to everyday interactions,”
2018
Later among the works it cites.
Georgia Tech, “Extended GTEA Gaze+,” http://webshare.ipat.gatech.edu/coc-rim-wall-lab/web/yli440/egtea_gp, 2018
2018
Later among the works it cites.
Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., Alahari, K.: Charades-ego: A large-scale dataset of paired third and first person videos. In: ArXiv (2018)
2018
Later among the works it cites.
L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,”
2018
Later among the works it cites.
B. Zhou, A. Andonian, A. Oliva, A. Torralba, “PyTorch Temporal Relational Networks,” https://github.com/metalbubble/TRN-pytorch, 2018
2018
Later among the works it cites.
J. Lin, C. Gan, and S. Han, “PyTorch Temporal Shift Module,” https://github.com/mit-han-lab/temporal-shift-module, 2018
2018
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” in
2019
Later among the works it cites.