Fetching the paper…
Reading the bibliography…
What will the future be? We wonder! In this survey, we explore the gap between current research in egocentric vision and the ever-anticipated future, where wearable computing, with outward facing cameras and digital overlays, is expected to be integrated in our every day lives.
Wolf W (1996) Key frame selection by motion analysis. In: ICASSP
1996
Earlier work this paper cites.
Pavlovic VI, Sharma R, Huang TS (1997) Visual Interpretation of Hand Gestures for Human-Computer Interaction: A Review. TPAMI 19(7):677–695
1997
Earlier work this paper cites.
Aoki H, Schiele B, Pentland A (1998) Recognizing Personal Location from Video. In: Workshop on Perceptual User Interfaces
1998
Earlier work this paper cites.
DeMenthon D, Kobla V, Doermann D (1998) Video summarization by curve simplification. In: International Conference on Multimedia
1998
Earlier work this paper cites.
Itti L, Koch C, Niebur E (1998) A Model of Saliency-Based Visual Attention for Rapid Scene Analysis. TPAMI 20(11):1254–1259
1998
Earlier work this paper cites.
Starner T, Schiele B, Pentland A (1998) Visual contextual awareness in wearable computing. In: International Symposium on Wearable Computers
1998
Earlier work this paper cites.
Farringdon J, Oni V (2000) Visual Augmented Memory (VAM). In: International Symposium on Wearable Computers
2000
Earlier work this paper cites.
Gers FA, Schmidhuber J, Cummins F (2000) Learning to Forget: Continual Prediction with LSTM. Neural computation 12(10):2451–2471
2000
Earlier work this paper cites.
Aizawa K, Ishijima K, Shiina M (2001) Summarizing wearable video. In: ICIP
2001
Earlier work this paper cites.
Davison AJ (2003) Real-time simultaneous localisation and mapping with a single camera. In: ICCV
2003
Earlier work this paper cites.
Torralba A, Murphy KP, Freeman WT, Rubin MA (2003) Context-based vision system for place and object recognition. In: ICCV
2003
Earlier work this paper cites.
Johnson M, Demiris Y (2005) Perceptual Perspective Taking and Action Recognition. International Journal of Advanced Robotic Systems 2(4):32
2005
Earlier work this paper cites.
Krishna S, Little G, Black J, Panchanathan S (2005) A wearable face recognition system for individuals with visual impairments. In: International Conference on Computers and Accessibility
2005
Earlier work this paper cites.
Mayol WW, Davison AJ, Tordoff BJ, Murray DW (2005) Applying Active Vision and SLAM to Wearables. In: Robotics Research
2005
Earlier work this paper cites.
Surie D, Pederson T, Lagriffoul F, Janlert LE, Sjölie D (2007) Activity Recognition Using an Egocentric Perspective of Everyday Objects. In: International Conference on Ubiquitous Intelligence and Computing
2007
Earlier work this paper cites.
Firat AK, Woon WL, Madnick S (2008) Technological Forecasting - A Review. Composite Information Systems Laboratory (CISL), Massachusetts Institute of Technology pp 1–19
2008
Earlier work this paper cites.
Irschara A, Zach C, Frahm JM, Bischof H (2009) From structure-from-motion point clouds to fast location recognition. In: CVPR
2009
Earlier work this paper cites.
Spriggs EH, De La Torre F, Hebert M (2009) Temporal segmentation and activity classification from first-person sensing. In: CVPR Workshop
2009
Earlier work this paper cites.
Castle RO, Klein G, Murray DW (2010) Combining monoSLAM with object recognition for scene augmentation using a wearable camera. Image and Vision Computing 28(11):1548–1556
2010
Earlier work this paper cites.
Jégou H, Douze M, Schmid C, Pérez P (2010) Aggregating local descriptors into a compact image representation. In: CVPR
2010
Earlier work this paper cites.
Ren X, Gu C (2010) Figure-ground segmentation improves handled object recognition in egocentric video. In: CVPR
2010
Earlier work this paper cites.
Badino H, Kanade T (2011) A Head-Wearable Short-Baseline Stereo System for the Simultaneous Estimation of Structure and Motion. In: International Conference on Machine Vision Applications
2011
Earlier work this paper cites.
Fathi A, Ren X, Rehg JM (2011) Learning to recognize objects in egocentric activities. In: CVPR
2011
Earlier work this paper cites.
Kang H, Hebert M, Kanade T (2011) Discovering object instances from scenes of Daily Living. In: ICCV
2011
Earlier work this paper cites.
Kitani KM, Okabe T, Sato Y, Sugimoto A (2011) Fast unsupervised ego-action learning for first-person sports videos. In: CVPR
2011
Earlier work this paper cites.
Kurze M, Roselius A (2011) Smart glasses linking real live and social network’s contacts by face recognition. In: Augmented Humans International Conference
2011
Earlier work this paper cites.
Oikonomidis I, Kyriazis N, Argyros AA (2011) Efficient model-based 3D tracking of hand articulations using Kinect. In: BMVC
2011
Earlier work this paper cites.
Pei M, Jia Y, Zhu SC (2011) Parsing video events with goal inference and intent prediction. In: ICCV
2011
Earlier work this paper cites.
Sattler T, Leibe B, Kobbelt L (2011) Fast image-based localization using direct 2D-to-3D matching. In: ICCV
2011
Earlier work this paper cites.
Shiratori T, Park HS, Sigal L, Sheikh Y, Hodgins JK (2011) Motion capture from body-mounted cameras. Transactions on Graphics 30(4):1–10
2011
Earlier work this paper cites.
Yamada K, Sugano Y, Okabe T, Sato Y, Sugimoto A, Hiraki K (2011) Can Saliency Map Models Predict Human Egocentric Visual Attention? In: ACCV Workshop
2011
Earlier work this paper cites.
Alcantarilla PF, Yebes JJ, Almazán J, Bergasa LM (2012) On combining visual SLAM and dense scene flow to increase the robustness of localization and mapping in dynamic environments. In: ICRA
2012
Earlier work this paper cites.
Gálvez-López D, Tardos JD (2012) Bags of Binary Words for Fast Place Recognition in Image Sequences. Transactions on Robotics 28(5):1188–1197
2012
Earlier work this paper cites.
Keskin C, Kıraç F, Kara YE, Akarun L (2012) Hand Pose Estimation and Hand Shape Classification Using Multi-layered Randomized Decision Forests. In: ECCV
2012
Earlier work this paper cites.
Lee YJ, Ghosh J, Grauman K (2012) Discovering important people and objects for egocentric video summarization. In: CVPR
2012
Earlier work this paper cites.
Murillo AC, Gutiérrez-Gómez D, Rituerto A, Puig L, Guerrero JJ (2012) Wearable omnidirectional vision system for personal localization and guidance. In: CVPR Workshop
2012
Earlier work this paper cites.
Park H, Jain E, Sheikh Y (2012) 3D Social Saliency from Head-mounted Cameras. In: NeurIPS
2012
Earlier work this paper cites.
Pirsiavash H, Ramanan D (2012) Detecting activities of daily living in first-person camera views. In: CVPR
2012
Earlier work this paper cites.
Shiraga K, Trung NT, Mitsugami I, Mukaigawa Y, Yagi Y (2012) Gait-based person authentication by wearable cameras. In: International Conference on Networked Sensing Systems
2012
Earlier work this paper cites.
Templeman R, Rahman Z, Crandall DJ, Kapadia A (2012) PlaceRaider: Virtual Theft in Physical Spaces with Smartphones. arXiv preprint arXiv:1209598
2012
Earlier work this paper cites.
Yamada K, Sugano Y, Okabe T, Sato Y, Sugimoto A, Hiraki K (2012) Attention Prediction in Egocentric Video Using Motion and Visual Saliency. In: Pacific-Rim Symposium on Image and Video Technology
2012
Earlier work this paper cites.
Ye Z, Li Y, Fathi A, Han Y, Rozga A, Abowd GD, Rehg JM (2012) Detecting eye contact using wearable eye-tracking glasses. In: International Joint Conference on Pervasive and Ubiquitous Computing
2012
Earlier work this paper cites.
Ionescu C, Papava D, Olaru V, Sminchisescu C (2013) Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments. TPAMI 36(7):1325–1339
2013
Earlier work this paper cites.
Khosla A, Hamid R, Lin CJ, Sundaresan N (2013) Large-Scale Video Summarization Using Web-Image Priors. In: CVPR
2013
Earlier work this paper cites.
Li Y, Fathi A, Rehg JM (2013) Learning to Predict Gaze in Egocentric Video. In: ICCV
2013
Earlier work this paper cites.
Lu Z, Grauman K (2013) Story-Driven Summarization for Egocentric Video. In: CVPR
2013
Earlier work this paper cites.
Park HS, Jain E, Sheikh Y (2013) Predicting Primary Gaze Behavior Using Social Saliency Fields. In: ICCV
2013
Earlier work this paper cites.
Ryoo MS, Matthies L (2013) First-Person Activity Recognition: What Are They Doing to Me? In: ICPR
2013
Earlier work this paper cites.
Smith BA, Yin Q, Feiner SK, Nayar SK (2013) Gaze locking: passive eye contact detection for human-object interaction. In: Symposium on User Interface Software and Technology
2013
Earlier work this paper cites.
Tang D, Yu TH, Kim TK (2013) Real-Time Articulated Hand Pose Estimation Using Semi-supervised Transductive Regression Forests. In: ICCV
2013
Earlier work this paper cites.
Thomaz E, Parnami A, Bidwell J, Essa I, Abowd GD (2013) Technological approaches for addressing privacy concerns when recognizing eating behaviors with wearable cameras. In: International Joint Conference on Pervasive and Ubiquitous Computing
2013
Earlier work this paper cites.
Wang X, Zhao X, Prakash V, Shi W, Gnawali O (2013) Computerized-Eyewear Based Face Recognition System for Improving Social Lives of Prosopagnosics. In: International Conference on Pervasive Computing Technologies for Healthcare
2013
Earlier work this paper cites.
Arev I, Park HS, Sheikh Y, Hodgins J, Shamir A (2014) Automatic editing of footage from multiple social cameras. Transactions on Graphics 33(4):1–11
2014
Earlier work this paper cites.
Baraldi L, Paci F, Serra G, Benini L, Cucchiara R (2014) Gesture Recognition in Ego-centric Videos Using Dense Trajectories and Hand Segmentation. In: CVPR Workshop
2014
Earlier work this paper cites.
Capi G, Kitani M, Ueki K (2014) Guide robot intelligent navigation in urban environments. Advanced Robotics 28(15):1043–1053
2014
Earlier work this paper cites.
Damen D, Leelasawassuk T, Haines O, Calway A, Mayol-Cuevas W (2014) You-Do, I-Learn: Discovering Task Relevant Objects and their Modes of Interaction from Multi-User Egocentric Video. In: BMVC
2014
Earlier work this paper cites.
Denning T, Dehlawi Z, Kohno T (2014) In situ with bystanders of augmented reality glasses: perspectives on recording and privacy-mediating technologies. In: Conference on Human Factors in Computing Systems
2014
Earlier work this paper cites.
Dey A, Billinghurst M, Lindeman RW, Swan JE (2018) A Systematic Review of 10 Years of Augmented Reality Usability Studies: 2005 to 2014. Frontiers in Robotics and AI 5:37
2014
Earlier work this paper cites.
Gygli M, Grabner H, Riemenschneider H, Van Gool L (2014) Creating Summaries from User Videos. In: ECCV
2014
Earlier work this paper cites.
Hoshen Y, Ben-Artzi G, Peleg S (2014) Wisdom of the Crowd in Egocentric Video Curation. In: CVPR Workshop
2014
Earlier work this paper cites.
Hoyle R, Templeman R, Armes S, Anthony D, Crandall D, Kapadia A (2014) Privacy behaviors of lifeloggers using wearable cameras. In: International Joint Conference on Pervasive and Ubiquitous Computing
2014
Earlier work this paper cites.
Kopf J, Cohen MF, Szeliski R (2014) First-person hyper-lapse videos. Transactions on Graphics 33(4):1–10
2014
Earlier work this paper cites.
Lan T, Chen TC, Savarese S (2014) A Hierarchical Representation for Future Action Prediction. In: ECCV
2014
Earlier work this paper cites.
Narayan S, Kankanhalli MS, Ramakrishnan KR (2014) Action and Interaction Recognition in First-Person Videos. In: CVPR Workshop
2014
Earlier work this paper cites.
Okamoto M, Yanai K (2014) Summarization of Egocentric Moving Videos for Generating Walking Route Guidance. In: Pacific-Rim Symposium on Image and Video Technology
2014
Earlier work this paper cites.
Petric F, Hrvatinić K, Babić A, Malovan L, Miklić D, Kovačić Z, Cepanec M, Stošić J, Šimleša S (2014) Four tasks of a robot-assisted autism spectrum disorder diagnostic protocol: First clinical tests. In: Global Humanitarian Technology Conference
2014
Earlier work this paper cites.
Qian C, Sun X, Wei Y, Tang X, Sun J (2014) Realtime and Robust Hand Tracking from Depth. In: CVPR
2014
Earlier work this paper cites.
Roesner F, Kohno T, Molnar D (2014) Security and privacy for augmented reality systems. Communications of the ACM 57(4):88–96
2014
Earlier work this paper cites.
Rogez G, Khademi M, Supancic JS, Montiel JMM, Ramanan D (2014) 3D Hand Pose Detection in Egocentric RGB-D Images. In: ECCV Workshop
2014
Earlier work this paper cites.
Templeman R, Korayem M, Crandall DJ, Kapadia A (2014) PlaceAvoider: Steering First-Person Cameras away from Sensitive Spaces. In: Network and Distributed System Security Symposium
2014
Earlier work this paper cites.
Xiong B, Grauman K (2014) Detecting Snap Points in Egocentric Video with a Web Photo Prior. In: ECCV
2014
Earlier work this paper cites.
Zhao B, Xing EP (2014) Quasi Real-Time Summarization for Consumer Videos. In: CVPR
2014
Earlier work this paper cites.
Bambach S, Lee S, Crandall DJ, Yu C (2015) Lending A Hand: Detecting Hands and Recognizing Activities in Complex Egocentric Interactions. In: ICCV
2015
Earlier work this paper cites.
Bertasius G, Park HS, Shi J (2015) Exploiting Egocentric Object Prior for 3D Saliency Detection. arXiv preprint arXiv:151102682
2015
Earlier work this paper cites.
Betancourt A, Morerio P, Regazzoni CS, Rauterberg M (2015) The Evolution of First Person Vision Methods: A Survey. Transactions on Circuits and Systems for Video Technology 25(5):744–760
2015
Earlier work this paper cites.
Bolaños M, Radeva P (2015) Ego-object discovery. arXiv preprint arXiv:150401639
2015
Earlier work this paper cites.
Hoyle R, Templeman R, Anthony D, Crandall D, Kapadia A (2015) Sensitive Lifelogs: A Privacy Analysis of Photos from Wearable Cameras. In: Conference on Human Factors in Computing Systems
2015
Earlier work this paper cites.
Kendall A, Grimes M, Cipolla R (2015) PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization. In: ICCV
2015
Earlier work this paper cites.
Koppula HS, Saxena A (2015) Anticipating Human Activities Using Object Affordances for Reactive Robotic Response. TPAMI 38(1):14–29
2015
Earlier work this paper cites.
Kumano S, Otsuka K, Ishii R, Yamato J (2015) Automatic gaze analysis in multiparty conversations based on Collective First-Person Vision. In: International Conference and Workshops on Automatic Face and Gesture Recognition
2015
Earlier work this paper cites.
Li Y, Ye Z, Rehg JM (2015) Delving into egocentric actions. In: CVPR
2015
Earlier work this paper cites.
Lin Y, Abdelfatah K, Zhou Y, Fan X, Yu H, Qian H, Wang S (2015) Co-Interest Person Detection from Multiple Wearable Camera Videos. In: ICCV
2015
Earlier work this paper cites.
Mandal B, Chia SC, Li L, Chandrasekhar V, Tan C, Lim JH (2015) A Wearable Face Recognition System on Google Glass for Assisting Social Interactions. In: ACCV
2015
Earlier work this paper cites.
Park HS, Shi J (2015) Social saliency prediction. In: CVPR
2015
Earlier work this paper cites.
Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M, Berg AC, Fei-Fei L (2015) ImageNet Large Scale Visual Recognition Challenge. IJCV 115(3):211–252
2015
Earlier work this paper cites.
Ryoo MS, Rothrock B, Matthies L (2015) Pooled motion features for first-person videos. In: CVPR
2015
Earlier work this paper cites.
Song Y, Vallmitjana J, Stent A, Jaimes A (2015) TVSum: Summarizing web videos using titles. In: CVPR
2015
Earlier work this paper cites.
Xia L, Gori I, Aggarwal JK, Ryoo MS (2015) Robot-centric Activity Recognition from First-Person RGB-D Videos. In: WACV
2015
Earlier work this paper cites.
Xiong B, Kim G, Sigal L (2015) Storyline Representation of Egocentric Videos with an Applications to Story-Based Search. In: ICCV
2015
Earlier work this paper cites.
Xu J, Mukherjee L, Li Y, Warner J, Rehg JM, Singh V (2015) Gaze-enabled egocentric video summarization via constrained submodular maximization. In: CVPR
2015
Earlier work this paper cites.
Ye Z, Li Y, Liu Y, Bridges C, Rozga A, Rehg JM (2015) Detecting bids for eye contact using a wearable camera. In: International Conference and Workshops on Automatic Face and Gesture Recognition
2015
Earlier work this paper cites.
Yonetani R, Kitani KM, Sato Y (2015) Ego-surfing first person videos. In: CVPR
2015
Earlier work this paper cites.
Yu F, Seff A, Zhang Y, Song S, Funkhouser T, Xiao J (2015) LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop. arXiv preprint arXiv:150603365
2015
Earlier work this paper cites.
Zheng L, Shen L, Tian L, Wang S, Wang J, Tian Q (2015) Scalable Person Re-identification: A Benchmark. In: ICCV
2015
Earlier work this paper cites.
Ahmetovic D, Gleason C, Ruan C, Kitani K, Takagi H, Asakawa C (2016) NavCog: a navigational cognitive assistant for the blind. In: International Conference on Human-Computer Interaction with Mobile Devices and Services
2016
Earlier work this paper cites.
Arandjelovic R, Gronat P, Torii A, Pajdla T, Sivic J (2016) NetVLAD: CNN Architecture for Weakly Supervised Place Recognition. In: CVPR
2016
Earlier work this paper cites.
Ardeshir S, Borji A (2016) Ego2Top: Matching Viewers in Egocentric and Top-View Videos. In: ECCV
2016
Earlier work this paper cites.
Bettadapura V, Castro D, Essa I (2016) Discovering picturesque highlights from egocentric vacation videos. In: WACV
2016
Earlier work this paper cites.
Bolaños M, Dimiccoli M, Radeva P (2016) Toward Storytelling From Visual Lifelogging: An Overview. Transactions on Human-Machine Systems 47(1):77–90
2016
Earlier work this paper cites.
Cai M, Kitani KM, Sato Y (2016) Understanding Hand-Object Manipulation with Grasp Types and Object Attributes. In: Robotics: Science and Systems
2016
Earlier work this paper cites.
Chakraborty A, Mandal B, Galoogahi HK (2016) Person re-identification using multiple first-person-views on wearable devices. In: WACV
2016
Earlier work this paper cites.
Chan CS, Chen SZ, Xie P, Chang CC, Sun M (2016) Recognition from Hand Cameras: A Revisit with Deep Learning. In: ECCV
2016
Earlier work this paper cites.
Damen D, Leelasawassuk T, Mayol-Cuevas W (2016) You-Do, I-Learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance. CVIU 149:98–112
2016
Earlier work this paper cites.
De Smedt Q, Wannous H, Vandeborre JP (2016) Skeleton-Based Dynamic Hand Gesture Recognition. In: CVPR Workshop
2016
Earlier work this paper cites.
Del Molino AG, Tan C, Lim JH, Tan AH (2016) Summarization of Egocentric Videos: A Comprehensive Survey. Transactions on Human-Machine Systems 47(1):65–76
2016
Earlier work this paper cites.
Fergnani F, Alletto S, Serra G, De Mira J, Cucchiara R (2016) Body Part Based Re-Identification from an Egocentric Perspective. In: CVPR Workshop
2016
Earlier work this paper cites.
Furnari A, Farinella GM, Battiato S (2016) Temporal Segmentation of Egocentric Videos to Highlight Personal Locations of Interest. In: ECCV Workshop
2016
Earlier work this paper cites.
Gori I, Aggarwal J, Matthies L, Ryoo MS (2016) Multitype Activity Recognition in Robot-Centric Scenarios. Robotics and Automation Letters 1(1):593–600
2016
Earlier work this paper cites.
Gutierrez-Gomez D, Guerrero J (2016) True scaled 6 DoF egocentric localisation with monocular wearable systems. Image and Vision Computing 52:178–194
2016
Earlier work this paper cites.
Hoshen Y, Peleg S (2016) An Egocentric Look at Video Photographer Identity. In: CVPR
2016
Earlier work this paper cites.
Huang Y, Liu X, Zhang X, Jin L (2016) A Pointing Gesture Based Egocentric Interaction System: Dataset, Approach and Application. In: CVPR Workshop
2016
Earlier work this paper cites.
Kera H, Yonetani R, Higuchi K, Sato Y (2016) Discovering Objects of Joint Attention via First-Person Sensing. In: CVPR Workshop
2016
Earlier work this paper cites.
Korayem M, Templeman R, Chen D, Crandall D, Kapadia A (2016) Enhancing Lifelogging Privacy by Detecting Screens. In: Conference on Human Factors in Computing Systems
2016
Earlier work this paper cites.
Molchanov P, Yang X, Gupta S, Kim K, Tyree S, Kautz J (2016) Online Detection and Classification of Dynamic Hand Gestures with Recurrent 3D Convolutional Neural Networks. In: CVPR
2016
Earlier work this paper cites.
Nguyen THC, Nebel JC, Florez-Revuelta F (2016) Recognition of Activities of Daily Living with Egocentric Vision: A Review. Sensors 16(1):72
2016
Earlier work this paper cites.
Park HS, Hwang JJ, Niu Y, Shi J (2016) Egocentric Future Localization. In: CVPR
2016
Earlier work this paper cites.
Poleg Y, Ephrat A, Peleg S, Arora C (2016) Compact CNN for indexing egocentric videos. In: WACV
2016
Earlier work this paper cites.
Rhinehart N, Kitani KM (2016) Learning Action Maps of Large Environments via First-Person Vision. In: CVPR
2016
Earlier work this paper cites.
Rhodin H, Richardt C, Casas D, Insafutdinov E, Shafiei M, Seidel HP, Schiele B, Theobalt C (2016) EgoCap: egocentric marker-less motion capture with two fisheye cameras. Transactions on Graphics 35(6):1–11
2016
Earlier work this paper cites.
Ryoo MS, Rothrock B, Fleming C, Yang HJ (2016) Privacy-Preserving Human Activity Recognition from Extreme Low Resolution. In: Conference on Artificial Intelligence
2016
Earlier work this paper cites.
Sattler T, Leibe B, Kobbelt L (2016) Efficient & Effective Prioritized Matching for Large-Scale Image-Based Localization. TPAMI 39(9):1744–1756
2016
Earlier work this paper cites.
Sharghi A, Gong B, Shah M (2016) Query-Focused Extractive Video Summarization. In: ECCV
2016
Earlier work this paper cites.
Song S, Chandrasekhar V, Mandal B, Li L, Lim JH, Babu GS, San PP, Cheung NM (2016) Multimodal Multi-Stream Deep Learning for Egocentric Activity Recognition. In: CVPR Workshop
2016
Earlier work this paper cites.
Su S, Hong JP, Shi J, Park HS (2016) Social Behavior Prediction from First Person Videos. arXiv preprint arXiv:161109464
2016
Earlier work this paper cites.
Su YC, Grauman K (2016) Detecting Engagement in Egocentric Video. In: ECCV
2016
Earlier work this paper cites.
Vondrick C, Pirsiavash H, Torralba A (2016) Anticipating Visual Representations from Unlabeled Video. In: CVPR
2016
Earlier work this paper cites.
Yang JA, Lee CH, Yang SW, Somayazulu VS, Chen YK, Chien SY (2016) Wearable social camera: Egocentric video summarization for social interaction. In: International Conference on Multimedia & Expo Workshop
2016
Cited alongside, same era.
Yao T, Mei T, Rui Y (2016) Highlight Detection with Pairwise Deep Ranking for First-Person Video Summarization. In: CVPR
2016
Cited alongside, same era.
Yonetani R, Kitani KM, Sato Y (2016) Recognizing Micro-Actions and Reactions from Paired Egocentric Videos. In: CVPR
2016
Cited alongside, same era.
Zhang K, Chao WL, Sha F, Grauman K (2016) Video Summarization with Long Short-Term Memory. In: ECCV
2016
Cited alongside, same era.
Aghaei M, Dimiccoli M, Ferrer CC, Radeva P (2017) Social Style Characterization from Egocentric Photo-streams. In: ICCV Workshop
2017
Cited alongside, same era.
Wieczorek M, Rychalska B, Dąbrowski J (2021) On the Unreasonable Effectiveness of Centroids in Image Retrieval. In: NeurIPS
2021
Later among the works it cites.
Yu X, Rao Y, Zhao W, Lu J, Zhou J (2021) Group-aware Contrastive Regression for Action Quality Assessment. In: ICCV
2021
Later among the works it cites.
Akada H, Wang J, Shimada S, Takahashi M, Theobalt C, Golyanik V (2022) UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture. In: ECCV
2022
Later among the works it cites.
Alikadic A, Saito H, Hachiuma R (2022) Transformer Networks for Future Person Localization in First-Person Videos. In: International Symposium on Visual Computing
2022
Later among the works it cites.
Bansal S, Arora C, Jawahar C (2022) My View is the Best View: Procedure Learning from Egocentric Videos. In: ECCV
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bertasius G, Shi J (2017) Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention. In: ICCV Workshop
2017
Cited alongside, same era.
Bertasius G, Soo Park H, Yu SX, Shi J (2017) Unsupervised Learning of Important Objects from First-Person Videos. In: ICCV
2017
Cited alongside, same era.
Cao C, Zhang Y, Wu Y, Lu H, Cheng J (2017) Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks with Spatiotemporal Transformer Modules. In: ICCV
2017
Cited alongside, same era.
Fan C, Lee J, Xu M, Kumar Singh K, Jae Lee Y, Crandall DJ, Ryoo MS (2017) Identifying First-Person Camera Wearers in Third-Person Videos. In: CVPR
2017
Cited alongside, same era.
Furnari A, Battiato S, Grauman K, Farinella GM (2017) Next-active-object prediction from egocentric videos. Journal of Visual Communication and Image Representation 49:401–411
2017
Cited alongside, same era.
Gao J, Yang Z, Nevatia R (2017) RED: Reinforced Encoder-Decoder Networks for Action Anticipation. In: BMVC
2017
Cited alongside, same era.
Garcia-Hernando G, Yuan S, Baek S, Kim TK (2017) First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations. In: CVPR
2017
Cited alongside, same era.
Bärmann L, Waibel A (2022) Where did I leave my keys? — Episodic-Memory-Based Question Answering on Egocentric Videos. In: CVPR Workshop
2022
Later among the works it cites.
Berton G, Masone C, Caputo B (2022) Rethinking Visual Geo-localization for Large-Scale Applications. In: CVPR
2022
Later among the works it cites.
Chandio Y, Bashir N, Anwar FM (2022) HoloSet - A Dataset for Visual-Inertial Pose Estimation in Extended Reality: Dataset. In: Conference on Embedded Networked Sensor Systems
2022
Later among the works it cites.
Chen C, Anjum S, Gurari D (2022) Grounding Answers for Visual Questions Asked by Visually Impaired People. In: CVPR
2022
Later among the works it cites.
Cheng J, Zhang L, Chen Q, Hu X, Cai J (2022) A review of visual SLAM methods for autonomous driving vehicles. Engineering Applications of Artificial Intelligence 114:104992
2022
Later among the works it cites.
Damen D, Doughty H, Farinella GM, Furnari A, Ma J, Kazakos E, Moltisanti D, Munro J, Perrett T, Price W, Wray M (2022) Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100. IJCV 130:33–55
2022
Later among the works it cites.
Darkhalil A, Shan D, Zhu B, Ma J, Kar A, Higgins R, Fidler S, Fouhey D, Damen D (2022) EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations. In: NeurIPS
2022
Later among the works it cites.
Datta S, Dharur S, Cartillier V, Desai R, Khanna M, Batra D, Parikh D (2022) Episodic Memory Question Answering. In: CVPR
2022
Later among the works it cites.
Devagiri JS, Paheding S, Niyaz Q, Yang X, Smith S (2022) Augmented Reality and Artificial Intelligence in industry: Trends, tools, and future challenges. Expert Systems with Applications 207
2022
Later among the works it cites.
Elfeki M, Wang L, Borji A (2022) Multi-stream dynamic video Summarization. In: WACV
2022
Later among the works it cites.
Gabeur V, Seo PH, Nagrani A, Sun C, Alahari K, Schmid C (2022) AVATAR: Unconstrained Audiovisual Speech Recognition. In: INTERSPEECH
2022
Later among the works it cites.
Girdhar R, Singh M, Ravi N, van der Maaten L, Joulin A, Misra I (2022) Omnivore: A Single Model for Many Visual Modalities. In: CVPR
2022
Later among the works it cites.
Grauman K, Westbury A, Byrne E, Chavis Z, Furnari A, Girdhar R, Hamburger J, Jiang H, Liu M, Liu X, et al (2022) Ego4D: Around the World in 3,000 Hours of Egocentric Video. In: CVPR
2022
Later among the works it cites.
Herzig R, Ben-Avraham E, Mangalam K, Bar A, Chechik G, Rohrbach A, Darrell T, Globerson A (2022) Object-Region Video Transformers. In: CVPR
2022
Later among the works it cites.
Jiang H, Murdock C, Ithapu VK (2022) Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization. In: CVPR
2022
Later among the works it cites.
Kazerouni IA, Fitzgerald L, Dooly G, Toal D (2022) A survey of state-of-the-art on visual SLAM. Expert Systems with Applications 205:117734
2022
Later among the works it cites.
Lai B, Liu M, Ryan F, Rehg J (2022) In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation. In: BMVC
2022
Later among the works it cites.
Leonardi R, Ragusa F, Furnari A, Farinella GM (2022) Egocentric Human-Object Interaction Detection Exploiting Synthetic Data. In: ICIAP
2022
Later among the works it cites.
Li Y, Cao Z, Liang A, Liang B, Chen L, Zhao H, Feng C (2022) Egocentric Prediction of Action Target in 3D. In: CVPR
2022
Later among the works it cites.
Lin KQ, Wang J, Soldan M, Wray M, Yan R, Xu Z, Gao D, Tu RC, Zhao W, Kong W, Cai C, HongFa W, Damen D, Ghanem B, Liu W, Shou MZ (2022) Egocentric Video-Language Pretraining. In: NeurIPS
2022
Later among the works it cites.
Lu H, Brimijoin WO (2022) Sound Source Selection Based on Head Movements in Natural Group Conversation. Trends in Hearing 26:23312165221097789
2022
Later among the works it cites.
Massiceti D, Anjum S, Gurari D (2022) VizWiz grand challenge workshop at CVPR 2022. In: SIGACCESS Accessibility and Computing
2022
Later among the works it cites.
Nair S, Rajeswaran A, Kumar V, Finn C, Gupta A (2022) R3M: A Universal Visual Representation for Robot Manipulation. In: CoRL
2022
Later among the works it cites.
Ng T, Kim HJ, Lee VT, DeTone D, Yang TY, Shen T, Ilg E, Balntas V, Mikolajczyk K, Sweeney C (2022) NinjaDesc: Content-Concealing Visual Descriptors via Adversarial Learning. In: CVPR
2022
Later among the works it cites.
Núñez-Marcos A, Azkune G, Arganda-Carreras I (2022) Egocentric Vision-based Action Recognition: A survey. Neurocomputing 472:175–197
2022
Later among the works it cites.
Panek V, Kukelova Z, Sattler T (2022) MeshLoc: Mesh-Based Visual Localization. In: ECCV
2022
Later among the works it cites.
Pathirana P, Senarath S, Meedeniya D, Jayarathna S (2022) Eye gaze estimation: A survey on deep learning-based approaches. Expert Systems with Applications 199:116894
2022
Later among the works it cites.
Plizzari C, Planamente M, Goletto G, Cannici M, Gusso E, Matteucci M, Caputo B (2022) E2(GO)MOTION: Motion Augmented Event Stream for Egocentric Action Recognition. In: CVPR
2022
Later among the works it cites.
Purushwalkam S, Morgado P, Gupta A (2022) The Challenges of Continuous Self-Supervised Learning. In: ECCV
2022
Later among the works it cites.
Radosavovic I, Xiao T, James S, Abbeel P, Malik J, Darrell T (2022) Real-World Robot Learning with Masked Visual Pre-training. In: CoRL
2022
Later among the works it cites.
Rodin I, Furnari A, Mavroeidis D, Farinella GM (2022) Untrimmed Action Anticipation. In: ICIAP
2022
Later among the works it cites.
Roy D, Fernando B (2022) Action anticipation using latent goal learning. In: WACV
2022
Later among the works it cites.
de Santana Correia A, Colombini EL (2022) Attention, please! A survey of neural attention models in deep learning. Artificial Intelligence Review 55(8):6037–6124
2022
Later among the works it cites.
Sarlin PE, Dusmanu M, Schönberger JL, Speciale P, Gruber L, Larsson V, Miksik O, Pollefeys M (2022) LaMAR: Benchmarking Localization and Mapping for Augmented Reality. In: ECCV
2022
Later among the works it cites.
Sener F, Chatterjee D, Shelepov D, He K, Singhania D, Wang R, Yao A (2022) Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities. In: CVPR
2022
Later among the works it cites.
Shaw K, Bahl S, Pathak D (2022) VideoDex: Learning Dexterity from Internet Videos. In: CoRL
2022
Later among the works it cites.
Tango K, Ohkawa T, Furuta R, Sato Y (2022) Background Mixup Data Augmentation for Hand and Object-in-Contact Detection. In: ECCV Workshop
2022
Later among the works it cites.
Wang J, Liu L, Xu W, Sarkar K, Luvizon D, Theobalt C (2022) Estimating Egocentric 3D Human Pose in the Wild with External Weak Supervision. In: CVPR
2022
Later among the works it cites.
Wen H, Liu Y, Huang J, Duan B, Yi L (2022) Point Primitive Transformer for Long-Term 4D Point Cloud Video Understanding. In: ECCV
2022
Later among the works it cites.
Wong B, Chen J, Wu Y, Lei SW, Mao D, Gao D, Shou MZ (2022) AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant. In: ECCV
2022
Later among the works it cites.
Xiong X, Arnab A, Nagrani A, Schmid C (2022) M&M Mix: A Multimodal Multiview Transformer Ensemble. arXiv preprint arXiv:220609852
2022
Later among the works it cites.
Yan S, Xiong X, Arnab A, Lu Z, Zhang M, Sun C, Schmid C (2022) Multiview Transformers for Video Recognition. In: CVPR
2022
Later among the works it cites.
Yang J, Bhalgat Y, Chang S, Porikli F, Kwak N (2022) Dynamic Iterative Refinement for Efficient 3D Hand Pose Estimation. In: WACV
2022
Later among the works it cites.
Zheng Y, Yang Y, Mo K, Li J, Yu T, Liu Y, Liu CK, Guibas LJ (2022) GIMO: Gaze-Informed Human Motion Prediction in Context. In: ECCV
2022
Later among the works it cites.
Zhu K, Guo H, Yan T, Zhu Y, Wang J, Tang M (2022) PASS: Part-Aware Self-Supervised Pre-Training for Person Re-Identification. In: ECCV
2022
Later among the works it cites.
Akiva P, Huang J, Liang KJ, Kovvuri R, Chen X, Feiszli M, Dana K, Hassner T (2023) Self-Supervised Object Detection from Egocentric Videos. In: ICCV
2023
Closest in time.
Ali-bey A, Chaib-draa B, Giguère P (2023) MixVPR: Feature Mixing for Visual Place Recognition. In: WACV
2023
Closest in time.
Bandini A, Zariffa J (2023) Analysis of the Hands in Egocentric Vision: A Survey. TPAMI 45(6):6846–6866
2023
Closest in time.
Bao W, Chen L, Zeng L, Li Z, Xu Y, Yuan J, Kong Y (2023) Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting. In: ICCV
2023
Closest in time.
Bock M, Kuehne H, Van Laerhoven K, Moeller M (2023) WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition. arXiv preprint arXiv:230405088
2023
Closest in time.
Chelani K, Sattler T, Kahl F, Kukelova Z (2023) Privacy-Preserving Representations are not Enough: Recovering Scene Content from Camera Poses. In: CVPR
2023
Closest in time.
Chen Z, Chen S, Schmid C, Laptev I (2023) gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction. In: CVPR
2023
Closest in time.
Dancette C, Whitehead S, Maheshwary R, Vedantam R, Scherer S, Chen X, Cord M, Rohrbach M (2023) Improving Selective Visual Question Answering by Learning from Your Peers. In: CVPR
2023
Closest in time.
Dargan S, Bansal S, Kumar M, Mittal A, Kumar K (2023) Augmented Reality: A Comprehensive Review. Archives of Computational Methods in Engineering 30(2):1057–1080
2023
Closest in time.
Deng A, Yang T, Chen C (2023) A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition. In: ICCV
2023
Closest in time.
Dunnhofer M, Furnari A, Farinella GM, Micheloni C (2023) Visual Object Tracking in First Person Vision. IJCV 131(1):259–283
2023
Closest in time.
Fan Z, Taheri O, Tzionas D, Kocabas M, Kaufmann M, Black MJ, Hilliges O (2023) ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. In: CVPR
2023
Closest in time.
Gao D, Zhou L, Ji L, Zhu L, Yang Y, Shou MZ (2023) MIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-Form Video Question Answering. In: CVPR
2023
Closest in time.
Ghosh S, Dhall A, Hayat M, Knibbe J, Ji Q (2023) Automatic Gaze Analysis: A Survey of Deep Learning Based Approaches. TPAMI
2023
Closest in time.
Girdhar R, El-Nouby A, Liu Z, Singh M, Alwala KV, Joulin A, Misra I (2023) ImageBind: One Embedding Space To Bind Them All. In: CVPR
2023
Closest in time.
Gong X, Mohan S, Dhingra N, Bazin JC, Li Y, Wang Z, Ranjan R (2023) MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition. In: CVPR
2023
Closest in time.
Grauman K, Westbury A, Torresani L, Kitani K, Malik J, Afouras T, Ashutosh K, Baiyya V, Bansal S, Boote B, et al (2023) Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives. arXiv preprint arXiv:231118259
2023
Closest in time.
Haitz D, Jutzi B, Ulrich M, Jäger M, Hübner P (2023) Combining HoloLens with Instant-NeRFs: Advanced Real-Time 3D Mobile Mapping. arXiv preprint arXiv:230414301
2023
Closest in time.
Hampali S, Hodan T, Tran L, Ma L, Keskin C, Lepetit V (2023) In-Hand 3D Object Scanning from an RGB Sequence. In: CVPR
2023
Closest in time.
Hatano M, Hachiuma R, Saito H (2023) Trajectory Prediction in First-Person Video: Utilizing a Pre-Trained Bird’s-Eye View Model. In: International Conference on Computer Vision Theory and Applications
2023
Closest in time.
He B, Wang J, Qiu J, Bui T, Shrivastava A, Wang Z (2023) Align and Attend: Multimodal Summarization with Dual Contrastive Losses. In: CVPR
2023
Closest in time.
Huh J, Chalk J, Kazakos E, Damen D, Zisserman A (2023) Epic-Sounds: A Large-Scale Dataset of Actions that Sound. In: ICASSP
2023
Closest in time.
Hung-Cuong N, Nguyen TH, Scherer R, Le VH (2023) YOLO Series for Human Hand Action Detection and Classification from Egocentric Videos. Sensors 23
2023
Closest in time.
Jiang H, Ramakrishnan SK, Grauman K (2023) Single-Stage Visual Query Localization in Egocentric Videos. In: NeurIPS
2023
Closest in time.
Kai C, Haihua Z, Dunbing T, Kun Z (2023) Future pedestrian location prediction in first-person videos for autonomous vehicles and social robots. Image and Vision Computing 134:104671
2023
Closest in time.
Karunratanakul K, Prokudin S, Hilliges O, Tang S (2023) HARP: Personalized Hand Reconstruction from a Monocular RGB Video. In: CVPR
2023
Closest in time.
Khirodkar R, Bansal A, Ma L, Newcombe R, Vo M, Kitani K (2023) EgoHumans: An Egocentric 3D Multi-Human Benchmark. In: ICCV
2023
Closest in time.
Kurita S, Katsura N, Onami E (2023) RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4D. In: ICCV
2023
Closest in time.
Lange MD, Eghbalzadeh H, Tan R, Iuzzolino ML, Meier F, Ridgeway K (2023) EgoAdapt: A multi-stream evaluation study of adaptation to real-world egocentric user video. arXiv preprint arXiv:230705784
2023
Closest in time.
Lee J, Sung M, Choi H, Kim TK (2023) Im2Hands: Learning Attentive Implicit Representation of Interacting Two-Hand Shapes. In: CVPR
2023
Closest in time.
Leonardi R, Ragusa F, Furnari A, Farinella GM (2023) Exploiting Multimodal Synthetic Data for Egocentric Human-Object Interaction Detection in an Industrial Scenario. arXiv preprint arXiv:230612152
2023
Closest in time.
Li J, Liu K, Wu J (2023) Ego-Body Pose Estimation via Ego-Head Pose Estimation. In: CVPR
2023
Closest in time.
Mai J, Hamdi A, Giancola S, Zhao C, Ghanem B (2023) EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries. In: ICCV
2023
Closest in time.
Majumder S, Jiang H, Moulon P, Henderson E, Calamia P, Grauman K, Ithapu VK (2023) Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations. In: CVPR
2023
Closest in time.
Mangalam K, Akshulakov R, Malik J (2023) EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding. In: NeurIPS
2023
Closest in time.
Mascaró EV, Ahn H, Lee D (2023) Intention-Conditioned Long-Term Human Egocentric Action Anticipation. In: WACV
2023
Closest in time.
Mur-Labadia L, Guerrero JJ, Martinez-Cantin R (2023) Multi-label affordance mapping from egocentric vision. In: ICCV
2023
Closest in time.
Nagarajan T, Ramakrishnan SK, Desai R, Hillis J, Grauman K (2023) EgoEnv: Human-centric environment representations from egocentric video. In: NeurIPS
2023
Closest in time.
Ohkawa T, He K, Sener F, Hodan T, Tran L, Keskin C (2023) AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose Estimation. In: CVPR
2023
Closest in time.
Pasca RG, Gavryushin A, Kuo YL, Hilliges O, Wang X (2023) Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction. arXiv preprint arXiv:230109209
2023
Closest in time.
Perrett T, Sinha S, Burghardt T, Mirmehdi M, Damen D (2023) Use Your Head: Improving Long-Tail Video Recognition. In: CVPR
2023
Closest in time.
Pietrantoni M, Humenberger M, Sattler T, Csurka G (2023) SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual Localization. In: CVPR
2023
Closest in time.
Plizzari C, Perrett T, Caputo B, Damen D (2023) What can a cook in Italy teach a mechanic in India? Action Recognition Generalisation Over Scenarios and Locations. In: ICCV
2023
Closest in time.
Pramanick S, Song Y, Nag S, Lin KQ, Shah H, Shou MZ, Chellappa R, Zhang P (2023) EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone. In: ICCV
2023
Closest in time.
Qian S, Fouhey DF (2023) Understanding 3D Object Interaction from a Single Image. In: ICCV
2023
Closest in time.
Qiu J, Lo FPW, Gu X, Jobarteh M, Jia W, Baranowski T, Steiner M, Anderson A, McCrory M, Sazonov E, Sun M, Frost G, Lo B (2023) Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring. Transactions on Cybernetics pp 1–14
2023
Closest in time.
Radevski G, Grujicic D, Blaschko M, Moens MF, Tuytelaars T (2023) Multimodal Distillation for Egocentric Action Recognition. In: ICCV
2023
Closest in time.
Ramakrishnan SK, Al-Halah Z, Grauman K (2023) NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory. In: CVPR
2023
Closest in time.
Ravi S, Climent-Perez P, Morales T, Huesca-Spairani C, Hashemifard K, Flórez-Revuelta F (2023) ODIN: An OmniDirectional INdoor dataset capturing Activities of Daily Living from multiple synchronized modalities. In: CVPR
2023
Closest in time.
Reza S, Sundareshan B, Moghaddam M, Camps OI (2023) Enhancing Transformer Backbone for Egocentric Video Action Segmentation. In: CVPR Workshop
2023
Closest in time.
Rosinol A, Leonard JJ, Carlone L (2023) NeRF-SLAM: Real-Time Dense Monocular SLAM with Neural Radiance Fields. In: IROS
2023
Closest in time.
Ryan F, Jiang H, Shukla A, Rehg JM, Ithapu VK (2023) Egocentric Auditory Attention Localization in Conversations. In: CVPR
2023
Closest in time.
Sarlin PE, DeTone D, Yang TY, Avetisyan A, Straub J, Malisiewicz T, Bulo SR, Newcombe R, Kontschieder P, Balntas V (2023) OrienterNet: Visual Localization in 2D Public Maps with Neural Matching. In: CVPR
2023
Closest in time.
Shah A, Lundell B, Sawhney H, Chellappa R (2023) STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos. In: ICCV
2023
Closest in time.
Shapovalov R, Kleiman Y, Rocco I, Novotny D, Vedaldi A, Chen C, Kokkinos F, Graham B, Neverova N (2023) Replay: Multi-modal Multi-view Acted Videos for Casual Holography. In: ICCV
2023
Closest in time.
Tan S, Nagarajan T, Grauman K (2023) EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding. In: NeurIPS
2023
Closest in time.
Tendulkar P, Surís D, Vondrick C (2023) FLEX: Full-Body Grasping Without Full-Body Grasps. In: CVPR
2023
Closest in time.
Tokmakov P, Li J, Gaidon A (2023) Breaking the “Object” in Video Object Segmentation. In: CVPR
2023
Closest in time.
Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, Bashlykov N, Batra S, Bhargava P, Bhosale S, et al (2023) Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:230709288
2023
Closest in time.
Tschernezki V, Darkhalil A, Zhu Z, Fouhey D, Larina I, Larlus D, Damen D, Vedaldi A (2023) EPIC Fields: Marrying 3D Geometry and Video Understanding. In: NeurIPS
2023
Closest in time.
Tse THE, Mueller F, Shen Z, Tang D, Beeler T, Dou M, Zhang Y, Petrovic S, Chang HJ, Taylor J, Doosti B (2023) Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp 14666–14677
2023
Closest in time.
Vahdani E, Tian Y (2023) Deep Learning-Based Action Detection in Untrimmed Videos: A Survey. TPAMI 45(4):4302–4320
2023
Closest in time.
Wu JZ, Zhang DJ, Hsu W, Zhang M, Shou MZ (2023) Label-Efficient Online Continual Object Detection in Streaming Video. In: ICCV
2023
Closest in time.
Xue Z, Grauman K (2023) Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment. In: NeurIPS
2023
Closest in time.
Xue Z, Song Y, Grauman K, Torresani L (2023) Egocentric Video Task Translation. In: CVPR
2023
Closest in time.
Yang X, Chu FJ, Feiszli M, Goyal R, Torresani L, Tran D (2023) Relational Space-Time Query in Long-Form Videos. In: CVPR
2023
Closest in time.
Yu J, Li X, Zhao X, Zhang H, Wang YX (2023) Video State-Changing Object Segmentation. In: ICCV
2023
Closest in time.
Zatsarynna O, Gall J (2023) Action Anticipation with Goal Consistency. In: ICIP
2023
Closest in time.
Zhong Z, Schneider D, Voit M, Stiefelhagen R, Beyerer J (2023) Anticipative Feature Fusion Transformer for Multi-Modal Action Anticipation. In: WACV
2023
Closest in time.
Zhou X, Arnab A, Sun C, Schmid C (2023) How can objects help action recognition? In: CVPR
2023
Closest in time.
Pavlakos G, Shan D, Radosavovic I, Kanazawa A, Fouhey D, Malik J (2024) Reconstructing hands in 3d with transformers. arXiv preprint arXiv:231205251
2024
Closest in time.
Roy D, Rajendiran R, Fernando B (2024) Interaction Region Visual Transformer for Egocentric Action Anticipation. In: WACV
2024
Closest in time.
Yang M, Du Y, Ghasemipour K, Tompson J, Schuurmans D, Abbeel P (2024) Learning Interactive Real-World Simulators. In: ICLR
2024
Closest in time.
Cipresso P, Giglioli IAC, Raya MA, Riva G (2018) The Past, Present, and Future of Virtual and Augmented Reality Research: A Network and Cluster Analysis of the Literature. Frontiers in psychology p 2086
2086
Closest in time.
Xu W, Chatterjee A, Zollhoefer M, Rhodin H, Fua P, Seidel HP, Theobalt C (2019) Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera. Transactions on Visualization and Computer Graphics 25(5):2093–2101
2093
Closest in time.