Fetching the paper…
Reading the bibliography…
Understanding human actions from videos of first-person view poses significant challenges.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff · 1978
Earlier work this paper cites.
Using dynamic time warping to find patterns in time series
Donald J Berndt and James Clifford · 1994
Earlier work this paper cites.
The mirror-neuron system
Giacomo Rizzolatti and Laila Craighero · 2004
Earlier work this paper cites.
What is modelled during observational learning?
Nicola J Hodges, A Mark Williams, Spencer J Hayes, and Gavin Breslin · 2007
Earlier work this paper cites.
Guide to the carnegie mellon university multimodal activity (cmu-mmac) database
Fernando De la Torre, Jessica Hodgins, Adam Bargteil, Xavier Martin, Justin Macey, Alex Collado, and Pep Beltran · 2009
Earlier work this paper cites.
Attention prediction in egocentric video using motion and visual saliency
Kentaro Yamada, Yusuke Sugano, Takahiro Okabe, Yoichi Sato, Akihiro Sugimoto, and Kazuo Hiraki · 2012
Earlier work this paper cites.
Learning to predict gaze in egocentric video
Yin Li, Alireza Fathi, and James M Rehg · 2013
Earlier work this paper cites.
Cross-view action recognition over heterogeneous feature spaces
Xinxiao Wu, Han Wang, Cuiwei Liu, and Yunde Jia · 2013
Earlier work this paper cites.
The evolution of first person vision methods: A survey
Alejandro Betancourt, Pietro Morerio, Carlo S Regazzoni, and Matthias Rauterberg · 2015
Earlier work this paper cites.
Delving into egocentric actions
Yin Li, Zhefan Ye, and James M Rehg · 2015
Earlier work this paper cites.
Ego2top: Matching viewers in egocentric and top-view videos
Shervin Ardeshir and Ali Borji · 2016
Earlier work this paper cites.
Identifying first-person camera wearers in third-person videos
Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar Singh, Yong Jae Lee, David J Crandall, and Michael S Ryoo · 2017
Earlier work this paper cites.
Seeing invisible poses: Estimating 3d body pose from egocentric video
Hao Jiang and Kristen Grauman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
An exocentric look at egocentric actions and vice versa
Shervin Ardeshir and Ali Borji · 2018
Earlier work this paper cites.
Summarizing first-person videos from third persons’ points of view
Hsuan-I Ho, Wei-Chen Chiu, and Yu-Chiang Frank Wang · 2018
Earlier work this paper cites.
Predicting gaze in egocentric video by learning task-dependent attention transition
Yifei Huang, Minjie Cai, Zhenqiang Li, and Yoichi Sato · 2018
Earlier work this paper cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Earlier work this paper cites.
Actor and observer: Joint modeling of first and third-person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Joint person segmentation and identification in synchronized first-and third-person videos
Mingze Xu, Chenyou Fan, Yuchen Wang, Michael S Ryoo, and David J Crandall · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso · 2018
Earlier work this paper cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Lsta: Long short-term attention for egocentric action recognition
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang · 2020
Cited alongside, same era.
Lemma: A multi-view dataset for le arning multi-agent multi-task activities
Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu, and Song-Chun Zhu · 2020
Cited alongside, same era.
Learning navigation subroutines from egocentric videos
Ashish Kumar, Saurabh Gupta, and Jitendra Malik · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Retrieval augmented classification for long-tail visual recognition
Alexander Long, Wei Yin, Thalaiyasingam Ajanthan, Vu Nguyen, Pulak Purkait, Ravi Garg, Alan Blair, Chunhua Shen, and Anton van den Hengel · 2022
Later among the works it cites.
Retrieval-augmented transformer for image captioning
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2022
Later among the works it cites.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Cited alongside, same era.
You2me: Inferring body pose in egocentric video via first and second person interactions
Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman · 2020
Cited alongside, same era.
Natural language processing with Python and spaCy: A practical introduction
Yuli Vasiliev · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Cited alongside, same era.
Video captioning based on both egocentric and exocentric views of robot vision for human-robot interaction
Soo-Han Kang and Ji-Hyeong Han · 2021
Cited alongside, same era.
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al · 2022
Later among the works it cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2022
Later among the works it cites.
Actionformer: Localizing moments of actions with transformers
Chen-Lin Zhang, Jianxin Wu, and Yin Li · 2022
Later among the works it cites.
Retrieval-based language models and applications
Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen · 2023
Later among the works it cites.
Hiervl: Learning hierarchical video-language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman · 2023
Later among the works it cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al · 2023
Later among the works it cites.
Whisperx: Time-accurate speech transcription of long-form audio
Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman · 2023
Later among the works it cites.
Retrieval-enhanced contrastive vision-text models
Ahmet Iscen, Mathilde Caron, Alireza Fathi, and Cordelia Schmid · 2023
Later among the works it cites.
In the eye of transformer: Global–local correlation for egocentric gaze estimation and beyond
Bolin Lai, Miao Liu, Fiona Ryan, and James M Rehg · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2023
Later among the works it cites.
Naq: Leveraging narrations as queries to supervise episodic memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, and Kristen Grauman · 2023
Later among the works it cites.
Smallcap: lightweight image captioning prompted with retrieval augmentation
Rita Ramos, Bruno Martins, Desmond Elliott, and Yova Kementchedjhieva · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment
Zihui Xue and Kristen Grauman · 2023
Later among the works it cites.
Retrieval-augmented multimodal language modeling
Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih · 2023
Later among the works it cites.
Helping hands: An object-aware ego-centric video recognition model
Chuhan Zhang, Ankush Gputa, and Andrew Zisserman · 2023
Later among the works it cites.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Later among the works it cites.