Fetching the paper…
Reading the bibliography…
First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Learning significant locations and predicting user movement with gps
D. Ashbrook and T. Starner · 2002
Earlier work this paper cites.
Activity zones for context-aware computing
K. Koile, K. Tollmar, D. Demirdjian, H. Shrobe, and T. Darrell · 2003
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
L. v. d. Maaten and G. Hinton · 2008
Earlier work this paper cites.
Deep learning from temporal coherence in video
H. Mobahi, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Recognizing indoor scenes
A. Quattoni and A. Torralba · 2009
Earlier work this paper cites.
What makes a chair a chair?
H. Grabner, J. Gall, and L. Van Gool · 2011
Earlier work this paper cites.
From 3d scene geometry to human workspace
A. Gupta, S. Satkin, A. A. Efros, and M. Hebert · 2011
Earlier work this paper cites.
Scene recognition and weakly supervised object localization with deformable part-based models
M. Pandey and S. Lazebnik · 2011
Earlier work this paper cites.
Scene semantics from long-term observation of people
V. Delaitre, D. F. Fouhey, I. Laptev, J. Sivic, A. Gupta, and A. A. Efros · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
H. Pirsiavash and D. Ramanan · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Earlier work this paper cites.
Learning to predict gaze in egocentric video
Y. Li, A. Fathi, and J. M. Rehg · 2013
Earlier work this paper cites.
Story-driven summarization for egocentric video
Z. Lu and K. Grauman · 2013
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
S. Stein and S. J. McKenna · 2013
Earlier work this paper cites.
People watching: Human actions as a cue for single view geometry
D. F. Fouhey, V. Delaitre, A. Gupta, A. A. Efros, I. Laptev, and J. Sivic · 2014
Earlier work this paper cites.
Physically grounded spatio-temporal object affordances
H. S. Koppula and A. Saxena · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. Arslan, and T. Serre · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Scenegrok: Inferring action maps in 3d environments
M. Savva, A. X. Chang, P. Hanrahan, M. Fisher, and M. Nießner · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Predicting important objects for egocentric video summarization
Y. J. Lee and K. Grauman · 2015
Earlier work this paper cites.
Personal object discovery in first-person videos
C. Lu, R. Liao, and J. Jia · 2015
Earlier work this paper cites.
Orb-slam: a versatile and accurate monocular slam system
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos · 2015
Earlier work this paper cites.
Temporal perception and prediction in ego-centric video
Y. Zhou and T. L. Berg · 2015
Earlier work this paper cites.
Understanding hand-object manipulation with grasp types and object attributes
M. Cai, K. M. Kitani, and Y. Sato · 2016
Cited alongside, same era.
You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance
D. Damen, T. Leelasawassuk, and W. Mayol-Cuevas · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
An egocentric look at video photographer identity
Y. Hoshen and S. Peleg · 2016
Cited alongside, same era.
Slow and steady feature analysis: higher order temporal coherence in video
D. Jayaraman and K. Grauman · 2016
Cited alongside, same era.
Going deeper into first-person activity recognition
M. Ma, H. Fan, and K. M. Kitani · 2016
Demo2vec: Reasoning object affordances from online videos
K. Fang, T.-L. Wu, D. Yang, S. Savarese, and J. J. Lim · 2018
Later among the works it cites.
Personal-location-based temporal segmentation of egocentric videos for lifelogging applications
A. Furnari, S. Battiato, and G. M. Farinella · 2018
Later among the works it cites.
Detecting and recognizing human-object interactions
G. Gkioxari, R. Girshick, P. Dollár, and K. He · 2018
Later among the works it cites.
Mapnet: An allocentric spatial memory for mapping environments
J. F. Henriques and A. Vedaldi · 2018
Later among the works it cites.
Predicting gaze in egocentric video by learning task-dependent attention transition
Y. Huang, M. Cai, Z. Li, and Y. Sato · 2018
Later among the works it cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Y. Li, M. Liu, and J. M. Rehg · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning action maps of large environments via first-person vision
N. Rhinehart and K. M. Kitani · 2016
Cited alongside, same era.
Egocentric future localization
H. Soo Park, J.-J. Hwang, Y. Niu, and J. Shi · 2016
Cited alongside, same era.
Visual motif discovery via first-person vision
R. Yonetani, K. M. Kitani, and Y. Sato · 2016
Cited alongside, same era.
Joint discovery of object states and manipulation actions
J.-B. Alayrac, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2017
Cited alongside, same era.
Next-active-object prediction from egocentric videos
A. Furnari, S. Battiato, K. Grauman, and G. M. Farinella · 2017
Cited alongside, same era.
Red: Reinforced encoder-decoder networks for action anticipation
J. Gao, Z. Yang, and R. Nevatia · 2017
Cited alongside, same era.
Later among the works it cites.
Attend and interact: Higher-order object interactions for video understanding
C.-Y. Ma, A. Kadav, I. Melvin, Z. Kira, G. AlRegib, and H. Peter Graf · 2018
Later among the works it cites.
Semi-parametric topological memory for navigation
N. Savinov, A. Dosovitskiy, and V. Koltun · 2018
Later among the works it cites.
Action anticipation with rbf kernelized feature mapping rnn
Y. Shi, B. Fernando, and R. Hartley · 2018
Later among the works it cites.
Charades-ego: A large-scale dataset of paired third and first person videos
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari · 2018
Later among the works it cites.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Later among the works it cites.
Scene memory transformer for embodied agents in long-horizon tasks
K. Fang, A. Toshev, L. Fei-Fei, and S. Savarese · 2019
Later among the works it cites.
What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention
A. Furnari and G. M. Farinella · 2019
Later among the works it cites.
Generative hybrid representations for activity forecasting with no-regret learning
J. Guan, Y. Yuan, K. M. Kitani, and N. Rhinehart · 2019
Later among the works it cites.
Timeception for complex action recognition
N. Hussein, E. Gavves, and A. W. Smeulders · 2019
Later among the works it cites.
Videograph: Recognizing minutes-long human activities in videos
N. Hussein, E. Gavves, and A. W. Smeulders · 2019
Later among the works it cites.
Time-conditioned action anticipation in one shot
Q. Ke, M. Fritz, and B. Schiele · 2019
Later among the works it cites.
Leveraging the present to anticipate the future in videos
A. Miech, I. Laptev, J. Sivic, H. Wang, L. Torresani, and D. Tran · 2019
Later among the works it cites.
Grounded human-object interaction hotspots from video
T. Nagarajan, C. Feichtenhofer, and K. Grauman · 2019
Later among the works it cites.
Anticipation and next action forecasting in video: an end-to-end model with memory
F. Pirri, L. Mauro, E. Alati, V. Ntouskos, M. Izadpanahkakhk, and E. Omrani · 2019
Later among the works it cites.
Lsta: Long short-term attention for egocentric action recognition
S. Sudhakaran, S. Escalera, and O. Lanz · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick · 2019
Later among the works it cites.
A structured model for action detection
Y. Zhang, P. Tokmakov, M. Hebert, and C. Schmid · 2019
Later among the works it cites.
Semantic understanding of scenes through the ade20k dataset
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba · 2019
Later among the works it cites.