Fetching the paper…
Reading the bibliography…
The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality.
Dynamic programming algorithm optimization for spoken word recognition
Hiroaki Sakoe and Seibi Chiba · 1978
Earlier work this paper cites.
Using dynamic time warping to find patterns in time series
Donald J Berndt and James Clifford · 1994
Earlier work this paper cites.
View-invariance in action recognition
Cen Rao and Mubarak Shah · 2001
Earlier work this paper cites.
View-invariant representation and recognition of actions
Cen Rao, Alper Yilmaz, and Mubarak Shah · 2002
Earlier work this paper cites.
Free viewpoint action recognition using motion history volumes
Daniel Weinland, Remi Ronfard, and Edmond Boyer · 2006
Earlier work this paper cites.
Learning to recognize activities from the wrong view point
Ali Farhadi and Mostafa Kamali Tabrizi · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Making action recognition robust to occlusions and viewpoint changes
Daniel Weinland, Mustafa Özuysal, and Pascal Fua · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Cross-view action recognition via view knowledge transfer
Jingen Liu, Mubarak Shah, Benjamin Kuipers, and Silvio Savarese · 2011
Earlier work this paper cites.
View-invariant action recognition based on artificial neural networks
Alexandros Iosifidis, Anastasios Tefas, and Ioannis Pitas · 2012
Earlier work this paper cites.
Latent multitask learning for view-invariant action recognition
Behrooz Mahasseni and Sinisa Todorovic · 2013
Earlier work this paper cites.
Cross-view action recognition over heterogeneous feature spaces
Xinxiao Wu, Han Wang, Cuiwei Liu, and Yunde Jia · 2013
Earlier work this paper cites.
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis · 2013
Earlier work this paper cites.
Cross-view action recognition via a continuous virtual path
Zhong Zhang, Chunheng Wang, Baihua Xiao, Wen Zhou, Shuang Liu, and Cunzhao Shi · 2013
Earlier work this paper cites.
Inferring unseen views of people
Chao-Yeh Chen and Kristen Grauman · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Applications of lp-norms and their smooth approximations for gradient based learning vector quantization
Mandy Lange, Dietlind Zühlke, Olaf Holz, Thomas Villmann, and Saxonia-Germany Mittweida · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Learning a non-linear knowledge transfer model for cross-view action recognition
Hossein Rahmani and Ajmal Mian · 2015
Earlier work this paper cites.
Action recognition in the presence of one egocentric and multiple static cameras
Bilge Soran, Ali Farhadi, and Linda Shapiro · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Unsupervised feature extraction by time-contrastive learning and nonlinear ica
Aapo Hyvarinen and Hiroshi Morioka · 2016
Earlier work this paper cites.
Slow and steady feature analysis: Higher order temporal coherence in video
D. Jayaraman and K. Grauman · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
3d action recognition from novel viewpoints
Hossein Rahmani and Ajmal Mian · 2016
Earlier work this paper cites.
Recognizing fine-grained and composite activities using hand-centric features and script data
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, and Bernt Schiele · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
Attention is all you need
Finegym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Later among the works it cites.
View-invariant probabilistic embedding for human pose
Jennifer J Sun, Jiaping Zhao, Liang-Chieh Chen, Florian Schroff, Hartwig Adam, and Ting Liu · 2020
Later among the works it cites.
The ikea asm dataset: Understanding people assembling furniture through actions, objects and pose
Yizhak Ben-Shabat, Xin Yu, Fatemeh Saleh, Dylan Campbell, Cristian Rodriguez-Opazo, Hongdong Li, and Stephen Gould · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Representation learning via global temporal alignment and cycle-consistency
Isma Hadji, Konstantinos G Derpanis, and Allan D Jepson · 2021
Later among the works it cites.
Learning by aligning videos in time
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
View adaptive recurrent neural networks for high performance human action recognition from skeleton data
Pengfei Zhang, Cuiling Lan, Junliang Xing, Wenjun Zeng, Jianru Xue, and Nanning Zheng · 2017
Cited alongside, same era.
An exocentric look at egocentric actions and vice versa
Shervin Ardeshir and Ali Borji · 2018
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Cited alongside, same era.
Co-training of audio and video representations from self-supervised temporal synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram N Syed, Andrey Konin, Zeeshan Zia, and Quoc-Huy Tran · 2021
Later among the works it cites.
H2o: Two hands manipulating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo, and Marc Pollefeys · 2021
Later among the works it cites.
Action shuffle alternating learning for unsupervised action segmentation
Jun Li and Sinisa Todorovic · 2021
Later among the works it cites.
Ego-exo: Transferring visual representations from third-person to first-person videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman · 2021
Later among the works it cites.
Shaping embodied agent behavior with activity-context priors from egocentric video
Tushar Nagarajan and Kristen Grauman · 2021
Later among the works it cites.
Recognizing actions in videos from unseen viewpoints
AJ Piergiovanni and Michael S Ryoo · 2021
Later among the works it cites.
Look at what i’m doing: Self-supervised spatial grounding of narrations in instructional videos
Reuben Tan, Bryan Plummer, Kate Saenko, Hailin Jin, and Bryan Russell · 2021
Later among the works it cites.
My view is the best view: Procedure learning from egocentric videos
Siddhant Bansal, Chetan Arora, and CV Jawahar · 2022
Later among the works it cites.
Frame-wise action representations for long videos via sequence contrastive learning
Minghao Chen, Fangyun Wei, Chong Li, and Deng Cai · 2022
Later among the works it cites.
Epic-kitchens visor benchmark: Video segmentations and object relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Context-aware sequence alignment using 4d skeletal augmentation
Taein Kwon, Bugra Tekin, Siyu Tang, and Marc Pollefeys · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani · 2022
Later among the works it cites.
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xiaolong Wang · 2022
Later among the works it cites.
Learning to align sequential actions in the wild
Weizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet, Pascal Fua, and Marc Pollefeys · 2022
Later among the works it cites.
Hand-object interaction reasoning
Jian Ma and Dima Damen · 2022
Later among the works it cites.
Self-supervised learning for videos: A survey
Madeline C. Schiappa, Yogesh S. Rawat, and Mubarak Shah · 2022
Later among the works it cites.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao · 2022
Later among the works it cites.
Tempclr: Temporal alignment representation with contrastive learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen, Xudong Lin, Guangxing Han, and Shih-Fu Chang · 2022
Later among the works it cites.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Yifeng Zhu, Abhishek Joshi, Peter Stone, and Yuke Zhu · 2022
Later among the works it cites.
Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2023
Closest in time.
Procedure-aware pretraining for instructional video understanding
Honglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese, and Juan Carlos Niebles · 2023
Closest in time.