Fetching the paper…
Reading the bibliography…
Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff · 1978
Earlier work this paper cites.
Human-computer interaction using eye-gaze input
Thomas E Hutchinson, K Preston White, Worthy N Martin, Kelly C Reichert, and Lisa A Frey · 1989
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 1996
Earlier work this paper cites.
The mirror-neuron system
Giacomo Rizzolatti and Laila Craighero · 2004
Earlier work this paper cites.
Nltk: the natural language toolkit
Steven Bird · 2006
Earlier work this paper cites.
What is modelled during observational learning?
Nicola J Hodges, A Mark Williams, Spencer J Hayes, and Gavin Breslin · 2007
Earlier work this paper cites.
Observational learning
Albert Bandura · 2008
Earlier work this paper cites.
Wearable augmented reality system using gaze interaction
Hyung Min Park, Seok Han Lee, and Jong Soo Choi · 2008
Earlier work this paper cites.
Guide to the carnegie mellon university multimodal activity (cmu-mmac) database
Fernando De la Torre, Jessica Hodgins, Adam Bargteil, Xavier Martin, Justin Macey, Alex Collado, and Pep Beltran · 2009
Earlier work this paper cites.
Who is the expert? analyzing gaze data to predict expertise level in collaborative applications
Yan Liu, Pei Yun Hsueh, Jennifer Lai, Mirweis Sangin, MarcAntoine Nüssli, and Pierre Dillenbourg · 2009
Earlier work this paper cites.
Multi-view video summarization
Yanwei Fu, Yanwen Guo, Yanshu Zhu, Feng Liu, Chuanming Song, and Zhi-Hua Zhou · 2010
Earlier work this paper cites.
Human-robot cooperation based on interaction learning
Stéphane Lallée, Eiichi Yoshida, Anthony Mallet, Francesco Nori, Lorenzo Natale, Giorgio Metta, Felix Warneken, and Peter Ford Dominey · 2010
Earlier work this paper cites.
Making action recognition robust to occlusions and viewpoint changes
Daniel Weinland, Mustafa Özuysal, and Pascal Fua · 2010
Earlier work this paper cites.
Psychomotor control in a virtual laparoscopic surgery training environment: gaze control parameters differentiate novices from experts
Mark Wilson, John Mcgrath, Samuel Vine, James Brewer, David Defriend, and Richard Masters · 2010
Earlier work this paper cites.
Combining eye gaze input with a brain–computer interface for touchless human–computer interaction
Thorsten O Zander, Matti Gaertner, Christian Kothe, and Roman Vilimek · 2010
Earlier work this paper cites.
Understanding observational learning: An interbehavioral approach
Mitch J Fryling, Cristin Johnston, and Linda J Hayes · 2011
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Learning to predict gaze in egocentric video
Yin Li, Alireza Fathi, and James M Rehg · 2013
Earlier work this paper cites.
Cross-view action recognition over heterogeneous feature spaces
Xinxiao Wu, Han Wang, Cuiwei Liu, and Yunde Jia · 2013
Earlier work this paper cites.
From actemes to action: A strongly supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and KG Derpanis · 2013
Earlier work this paper cites.
Learning view-invariant sparse representations for cross-view action recognition
Jingjing Zheng and Zhuolin Jiang · 2013
Earlier work this paper cites.
Pupil: An open source platform for pervasive eye tracking and mobile gaze-based interaction
Moritz Kassner, William Patera, and Andreas Bulling · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Egocentric field-of-view localization using first-person point-of-view devices
Vinay Bettadapura, Irfan Essa, and Caroline Pantofaru · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Using gaze patterns to predict task intent in collaboration
Chien-Ming Huang, Sean Andrist, Allison Sauppé, and Bilge Mutlu · 2015
Earlier work this paper cites.
Delving into egocentric actions
Yin Li, Zhefan Ye, and James M Rehg · 2015
Earlier work this paper cites.
Gaze-enabled egocentric video summarization via constrained submodular maximization
Jia Xu, Lopamudra Mukherjee, Yin Li, Jamieson Warner, James M Rehg, and Vikas Singh · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Understanding hand-object manipulation with grasp types and object attributes
Minjie Cai, Kris M Kitani, and Yoichi Sato · 2016
Earlier work this paper cites.
A dataset and benchmarks for segmentation and recognition of gestures in robotic surgery
Narges Ahmidi, Lingling Tao, Shahin Sefati, Yixin Gao, Colin Lea, Benjamin Bejar Haro, Luca Zappella, Sanjeev Khudanpur, René Vidal, and Gregory D Hager · 2017
Earlier work this paper cites.
Am i a baller? basketball performance assessment from first-person videos
Gedas Bertasius, Hyun Soo Park, Stella X Yu, and Jianbo Shi · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Learning robot activities from first-person human videos using convolutional future regression
Jangwon Lee and Michael S Ryoo · 2017
Earlier work this paper cites.
Learning to predict intent from gaze during robotic hand-eye coordination
Yosef Razin and Karen Feigh · 2017
Earlier work this paper cites.
An exocentric look at egocentric actions and vice versa
Shervin Ardeshir and Ali Borji · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Earlier work this paper cites.
Who’s better? who’s best? pairwise deep ranking for skill determination
Hazel Doughty, Dima Damen, and Walterio Mayol-Cuevas · 2018
Earlier work this paper cites.
Summarizing first-person videos from third persons’ points of view
Hsuan-I Ho, Wei-Chen Chiu, and Yu-Chiang Frank Wang · 2018
Earlier work this paper cites.
Predicting gaze in egocentric video by learning task-dependent attention transition
Yifei Huang, Minjie Cai, Zhenqiang Li, and Yoichi Sato · 2018
Earlier work this paper cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Earlier work this paper cites.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Learning to compare: Relation network for few-shot learning
Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales · 2018
Earlier work this paper cites.
Joint person segmentation and identification in synchronized first-and third-person videos
Mingze Xu, Chenyou Fan, Yuchen Wang, Michael S Ryoo, and David J Crandall · 2018
Earlier work this paper cites.
Disentangled sequential autoencoder
Li Yingzhen and Stephan Mandt · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
Tianhe Yu, Chelsea Finn, Annie Xie, Sudeep Dasari, Tianhao Zhang, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Temporal attentive alignment for large-scale video domain adaptation
Min-Hung Chen, Zsolt Kira, Ghassan AlRegib, Jaekwon Yoo, Ruxin Chen, and Jian Zheng · 2019
Earlier work this paper cites.
The pros and cons: Rank-aware temporal attention for skill determination in long videos
Hazel Doughty, Walterio Mayol-Cuevas, and Dima Damen · 2019
Earlier work this paper cites.
Ms-tcn: Multi-stage temporal convolutional network for action segmentation
Yazan Abu Farha and Jurgen Gall · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention
Antonino Furnari and Giovanni Maria Farinella · 2019
Cited alongside, same era.
Distinit: Learning video representations without a single labeled video
Rohit Girdhar, Du Tran, Lorenzo Torresani, and Deva Ramanan · 2019
Cited alongside, same era.
Epic-tent: An egocentric video dataset for camping tent assembly
Youngkyoon Jang, Brian Sullivan, Casimir Ludwig, Iain Gilchrist, Dima Damen, and Walterio Mayol-Cuevas · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Explainable embodied agents through social cues: a review
Sebastian Wallkötter, Silvia Tulli, Ginevra Castellano, Ana Paiva, and Mohamed Chetouani · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Group-aware contrastive regression for action quality assessment
Xumin Yu, Yongming Rao, Wenliang Zhao, Jiwen Lu, and Jie Zhou · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Cited alongside, same era.
Eyes are faster than hands: A soft wearable robot learns user intention from the egocentric view
Daekyum Kim, Brian Byunghyun Kang, Kyu Bum Kim, Hyungmin Choi, Jeesoo Ha, Kyu-Jin Cho, and Sungho Jo · 2019
Cited alongside, same era.
Habitat: A Platform for Embodied AI Research
Manolis Savva*, Abhishek Kadian*, Oleksandr Maksymets*, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
The jester dataset: A large-scale video dataset of human gestures
Joanna Materzynska, Guillaume Berger, Ingo Bax, and Roland Memisevic · 2019
Cited alongside, same era.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
What and how well you performed? a multitask learning approach to action quality assessment
Paritosh Parmar and Brendan Tran Morris · 2019
Cited alongside, same era.
Lsta: Long short-term attention for egocentric action recognition
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Cited alongside, same era.
My view is the best view: Procedure learning from egocentric videos
Siddhant Bansal, Chetan Arora, and C.V. Jawahar · 2022
Later among the works it cites.
Dual-head contrastive domain adaptation for video action recognition
Victor G Turrisi da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, and Elisa Ricci · 2022
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2022
Later among the works it cites.
Epic-kitchens visor benchmark: Video segmentations and object relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen · 2022
Later among the works it cites.
Human hands as probes for interactive object understanding
Mohit Goyal, Sahil Modi, Rishabh Goyal, and Saurabh Gupta · 2022
Later among the works it cites.
Ego4d: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, and Eugene Byrne et al · 2022
Later among the works it cites.
Compound prototype matching for few-shot action recognition
Yifei Huang, Lijin Yang, and Yoichi Sato · 2022
Later among the works it cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn · 2022
Later among the works it cites.
Egotaskqa: Understanding human tasks in egocentric videos
Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang · 2022
Later among the works it cites.
Machine learning for technical skill assessment in surgery: a systematic review
Kyle Lam, Junhong Chen, Zeyu Wang, Fahad M Iqbal, Ara Darzi, Benny Lo, Sanjay Purkayastha, and James M Kinross · 2022
Later among the works it cites.
Surgical skill assessment via video semantic aggregation
Zhenqiang Li, Lin Gu, Weimin Wang, Ryosuke Nakamura, and Yoichi Sato · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Head and eye egocentric gesture recognition for human-robot interaction using eyewear cameras
Javier Marina-Miranda and V Javier Traver · 2022
Later among the works it cites.
E2 (go) motion: Motion augmented event stream for egocentric action recognition
Chiara Plizzari, Mirco Planamente, Gabriele Goletto, Marco Cannici, Emanuele Gusso, Matteo Matteucci, and Barbara Caputo · 2022
Later among the works it cites.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
F. Sener, D. Chatterjee, D. Shelepov, K. He, D. Singhania, R. Wang, and A. Yao · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
Internvideo: General video foundation models via generative and discriminative learning
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al · 2022
Later among the works it cites.
Assistq: Affordance-centric question-driven task completion for egocentric assistant
Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, and Mike Zheng Shou · 2022
Later among the works it cites.
Incomplete multi-view domain adaptation via channel enhancement and knowledge transfer
Haifeng Xia, Pu Wang, and Zhengming Ding · 2022
Later among the works it cites.
Advancing high-resolution video-language representation with large-scale video transcriptions
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo · 2022
Later among the works it cites.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.
Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition
Lijin Yang, Yifei Huang, Yusuke Sugano, and Yoichi Sato · 2022
Later among the works it cites.
Gimo: Gaze-informed human motion prediction in context
Yang Zheng, Yanchao Yang, Kaichun Mo, Jiaman Li, Tao Yu, Yebin Liu, Karen Liu, and Leonidas J Guibas · 2022
Later among the works it cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al · 2023
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, et al · 2023
Later among the works it cites.
Weakly supervised temporal sentence grounding with uncertainty-guided self-training
Yifei Huang, Lijin Yang, and Yoichi Sato · 2023
Later among the works it cites.
In the eye of transformer: Global–local correlation for egocentric gaze estimation and beyond
Bolin Lai, Miao Liu, Fiona Ryan, and James M Rehg · 2023
Later among the works it cites.
Egocentric planning for scalable embodied task achievement
Xiaotian Liu, Hector Palacios, and Christian Muise · 2023
Later among the works it cites.
Copilot: Human-environment collision prediction and localization from egocentric videos
Boxiao Pan, Bokui Shen, Davis Rempe, Despoina Paschalidou, Kaichun Mo, Yanchao Yang, and Leonidas J Guibas · 2023
Later among the works it cites.
An outlook into the future of egocentric vision
Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi · 2023
Later among the works it cites.
Multimodal distillation for egocentric action recognition
Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens, and Tinne Tuytelaars · 2023
Later among the works it cites.
Project aria: A new tool for egocentric multi-modal ai research
Kiran Somasundaram, Jing Dong, Huixuan Tang, Julian Straub, Mingfei Yan, Michael Goesele, Jakob Julian Engel, Renzo De Nardi, and Richard Newcombe · 2023
Later among the works it cites.
Unsupervised video domain adaptation for action recognition: A disentanglement perspective
Pengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren, Jing Jiang, Xiang Yin, et al · 2023
Later among the works it cites.
Learning interaction regions and motion trajectories simultaneously from egocentric demonstration videos
Jianjia Xin, Lichun Wang, Kai Xu, Chao Yang, and Baocai Yin · 2023
Later among the works it cites.
Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment
Zihui Xue and Kristen Grauman · 2023
Later among the works it cites.
Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations
Takuma Yagi, Misaki Ohashi, Yifei Huang, Ryosuke Furuta, Shungo Adachi, Toutai Mitsuyama, and Yoichi Sato · 2023
Later among the works it cites.
Fine-grained affordance annotation for egocentric hand-object interaction videos
Zecheng Yu, Yifei Huang, Ryosuke Furuta, Takuma Yagi, Yusuke Goutsu, and Yoichi Sato · 2023
Later among the works it cites.
X-board: an egocentric adaptive ar assistant for perception in indoor environments
Zhenning Zhang, Zhigeng Pan, Weiqing Li, and Zhiyong Su · 2023
Later among the works it cites.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Later among the works it cites.
Video mamba suite: State space model as a versatile alternative for video understanding
Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang · 2024
Closest in time.
Ego4d goal-step: Toward hierarchical understanding of procedural activities
Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani · 2024
Closest in time.
Retrieval-augmented egocentric video captioning
Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie · 2024
Closest in time.
Video as the new language for real-world decision making
Sherry Yang, Jacob Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schuurmans · 2024
Closest in time.