Fetching the paper…
Reading the bibliography…
Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications.
A new mathematical formulation for strapdown inertial navigation
John E. Bortz · 1971
Earlier work this paper cites.
A comparison of pedestrian dead-reckoning algorithms using a low-cost mems imu
A.R. Jimenez, F. Seco, C. Prieto, and J. Guevara · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Activity recognition using cell phone accelerometers
Jennifer R. Kwapisz, Gary M. Weiss, and Samuel A. Moore · 2011
Earlier work this paper cites.
Walk detection and step counting on unconstrained smartphones
Agata Brajdic and Robert K. Harle · 2013
Earlier work this paper cites.
Hierarchical, multi-sensor based classification of daily life activities: Comparison with state-of-the-art algorithms using a benchmark dataset
Heike Leutheuser, Dominik Schuldhaus, and Bjoern M. Eskofier · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2014
Earlier work this paper cites.
A method of pedestrian dead reckoning for smartphones using frequency domain analysis on patterns of acceleration and angular velocity
Masakatsu Kourogi and Takeshi Kurata · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor
Chen Chen, Roozbeh Jafari, and Nasser Kehtarnavaz · 2015
Earlier work this paper cites.
Learning image representations tied to egomotion
D. Jayaraman and K. Grauman · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelović and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
On-manifold preintegration for real-time visual–inertial odometry
Christian Forster, Luca Carlone, Frank Dellaert, and Davide Scaramuzza · 2017
Earlier work this paper cites.
Who let the dogs out? modeling dog behavior from visual data
Kiana Ehsani, Hessam Bagherinezhad, Joseph Redmon, Roozbeh Mottaghi, and Ali Farhadi · 2018
Earlier work this paper cites.
Modality distillation with multiple stream networks for action recognition
Nuno C. Garcia, Pietro Morerio, and Vittorio Murino · 2018
Earlier work this paper cites.
Real-time human activity recognition from accelerometer data using convolutional neural networks
Andrey Ignatov · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Earlier work this paper cites.
Fixing weight decay regularization in adam, 2018
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Optical flow guided feature: A fast and robust motion representation for video action recognition
Shuyang Sun, Zhanghui Kuang, Lu Sheng, Wanli Ouyang, and Wei Zhang · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Ridi: Robust imu double integration
Hang Yan, Qi Shan, and Yasutaka Furukawa · 2018
Cited alongside, same era.
ECO: efficient convolutional network for online video understanding
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox · 2018
Cited alongside, same era.
Vision and acceleration modalities: Partners for recognizing complex activities
Alexander Diete, Timo Sztyler, and Heiner Stuckenschmidt · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Movinets: Mobile video networks for efficient video recognition
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong · 2021
Later among the works it cites.
3d-to-2d distillation for indoor scene parsing
Zhengzhe Liu, Xiaojuan Qi, and Chi-Wing Fu · 2021
Later among the works it cites.
Improved deep representation learning for human activity recognition using imu sensors
Niall Lyons, Avik Santra, and Ashutosh Pandey · 2021
Later among the works it cites.
Vision and inertial sensing fusion for human action recognition: A review
Sharmin Majumder and Nasser Kehtarnavaz · 2021
Later among the works it cites.
Adafuse: Adaptive temporal fusion network for efficient action recognition
Yue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri, Leonid Karlinsky, Kate Saenko, Aude Oliva, and Rogerio Feris · 2021
Later among the works it cites.
AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Self-supervised moving vehicle tracking with stereo sound
C. Gan, H. Zhao, P. Chen, D. Cox, and A. Torralba · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Cited alongside, same era.
Scsampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Cited alongside, same era.
Deep learning for sensor-based activity recognition: A survey
Jindong Wang, Yiqiang Chen, Shuji Hao, Xiaohui Peng, and Lisha Hu · 2019
Cited alongside, same era.
Person independent recognition of head gestures from parametrised and raw signals recorded from inertial measurement unit
Anna Borowska-Terka and Pawel Strumillo · 2020
Cited alongside, same era.
Denoising imu gyroscopes with deep learning for open-loop attitude estimation
M. Brossard, S. Bonnabel, and A. Barrau · 2020
Cited alongside, same era.
Rameswar Panda, Chun-Fu Chen, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, and Rogerio Feris · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and João F. Henriques · 2021
Later among the works it cites.
Attention-based sensor fusion for human activity recognition using imu signals
Wenjin Tao, Haodong Chen, Md. Moniruzzaman, Ming C. Leu, Zhaozheng Yi, and Ruwen Qin · 2021
Later among the works it cites.
Zero-shot learning for imu-based activity recognition using video embeddings
Catherine Tong, Jinchen Ge, and Nicholas D Lane · 2021
Later among the works it cites.
How you move your head tells what you do: Self-supervised video representation learning with egocentric cameras and imu sensors
Satoshi Tsutsui, Ruta Desai, and Karl Ridgeway · 2021
Later among the works it cites.
Adaptive focus for efficient video recognition
Yulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song, Yizeng Han, and Gao Huang · 2021
Later among the works it cites.
Rio: Rotation-equivariance supervised learning of robust inertial odometry
Xiya Cao, Caifa Zhou, Dandan Zeng, and Yongliang Wang · 2022
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2022
Later among the works it cites.
Episodic memory question answering
Samyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai, Mukul Khanna, Dhruv Batra, and Devi Parikh · 2022
Later among the works it cites.
Omnivore: A Single Model for Many Visual Modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Later among the works it cites.
Egocentric prediction of action target in 3d
Yiming Li, Ziang Cao, Andrew Liang, Benjamin Liang, Luoyao Chen, Hang Zhao, and Chen Feng · 2022
Later among the works it cites.
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xiaolong Wang · 2022
Later among the works it cites.
Joint coding of visual input and eye/head position in v1 of freely moving mice
Philip R. L. Parker, Elliott T. T. Abe, Emmalyn S. P. Leonard, Dylan M. Martins, and Cristopher M. Niell · 2022
Later among the works it cites.
E2(go)motion: Motion augmented event stream for egocentric action recognition
Chiara Plizzari, Mirco Planamente, Gabriele Goletto, Marco Cannici, Emanuele Gusso, Matteo Matteucci, and Barbara Caputo · 2022
Later among the works it cites.
Multi-scale deep feature learning for human activity recognition using wearable sensors
Yin Tang, Lei Zhang, Fuhong Min, and Jun He · 2022
Later among the works it cites.
Efficient video transformers with spatial-temporal token selection
Junke Wang, Xitong Yang, Hengduo Li, Liu Li, Zuxuan Wu, and Yu-Gang Jiang · 2022
Later among the works it cites.
Adafocus v2: End-to-end training of spatial dynamic networks for video recognition
Yulin Wang, Yang Yue, Yuanze Lin, Haojun Jiang, Zihang Lai, Victor Kulikov, Nikita Orlov, Humphrey Shi, and Gao Huang · 2022
Later among the works it cites.
Efficient deep visual and inertial odometry with adaptive visual modality selection
Mingyu Yang, Yu Chen, and Hun-Seok Kim · 2022
Later among the works it cites.