Fetching the paper…
Reading the bibliography…
Understanding videos that contain multiple modalities is crucial, especially in egocentric videos, where combining various sensory inputs significantly improves tasks like action recognition and moment localization.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Moddrop: adaptive multi-modal gesture recognition
Natalia Neverova, Christian Wolf, Graham Taylor, and Florian Nebout · 2015
Earlier work this paper cites.
Revisiting batch normalization for practical domain adaptation
Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu, and Xiaodi Hou · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Earlier work this paper cites.
Kitting in the wild through online domain adaptation
Massimiliano Mancini, Hakan Karaoguz, Elisa Ricci, Patric Jensfelt, and Barbara Caputo · 2018
Earlier work this paper cites.
Learning factorized multimodal representations
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2018
Earlier work this paper cites.
Learning with privileged information via adversarial discriminative modality distillation
Nuno C Garcia, Pietro Morerio, and Vittorio Murino · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
Audio feature generation for missing modality problem in video action recognition
Hu-Cheng Lee, Chih-Yu Lin, Pin-Chun Hsu, and Winston H Hsu · 2019
Earlier work this paper cites.
Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation
Jian Liang, Dapeng Hu, and Jiashi Feng · 2020
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell · 2020
Earlier work this paper cites.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Earlier work this paper cites.
Online continual learning with natural distribution shifts: An empirical study with visual data
Zhipeng Cai, Ozan Sener, and Vladlen Koltun · 2021
Earlier work this paper cites.
Improving multimodal fusion via mutual dependency maximisation
Pierre Colombo, Emile Chapuis, Matthieu Labeau, and Chloé Clavel · 2021
Cited alongside, same era.
With a little help from my temporal context: Multimodal egocentric action recognition
Evangelos Kazakos, Jaesung Huh, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2021
Cited alongside, same era.
Smil: Multimodal learning with severely missing modality
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng · 2021
Cited alongside, same era.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Revisiting test time adaptation under online evaluation
Motasem Alfarra, Hani Itani, Alejandro Pardo, Shyma Alhuwaider, Merey Ramazanova, Juan C Pérez, Zhipeng Cai, Matthias Müller, and Bernard Ghanem · 2023
Later among the works it cites.
Real-time evaluation in online continual learning: A new paradigm
Yasir Ghunaim, Adel Bibi, Kumail Alhamoud, Motasem Alfarra, Hasan Abed Al Kader Hammoud, Ameya Prabhu, Philip HS Torr, and Bernard Ghanem · 2023
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2023
Later among the works it cites.
Epic-sounds: A large-scale dataset of actions that sound
Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, and Andrew Zisserman · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinming Zhao, Ruichen Li, and Qin Jin · 2021
Cited alongside, same era.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2022
Cited alongside, same era.
Back to the source: Diffusion-driven test-time adaptation
Jin Gao, Jialing Zhang, Xihui Liu, Trevor Darrell, Evan Shelhamer, and Dequan Wang · 2022
Cited alongside, same era.
Omnivore: A single model for many visual modalities. 2022 ieee
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Cited alongside, same era.
Takeshi Kojima, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Cited alongside, same era.
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Cited alongside, same era.
Multimodal prompting with missing modalities for visual recognition
Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu, and Chen-Yu Lee · 2023
Later among the works it cites.
Towards stable test-time adaptation in dynamic wild world
Shuaicheng Niu14, Jiaxiang Wu, Yifan Zhang, Zhiquan Wen, Yaofo Chen, Peilin Zhao, and Mingkui Tan15 · 2023
Later among the works it cites.
Multimodal distillation for egocentric action recognition
Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens, and Tinne Tuytelaars · 2023
Later among the works it cites.
Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in egocentric videos
Merey Ramazanova, Victor Escorcia, Fabian Caba, Chen Zhao, and Bernard Ghanem · 2023
Later among the works it cites.
Centre stage: Centricity-based audio-visual temporal action detection
Hanyuan Wang, Majid Mirmehdi, Dima Damen, and Toby Perrett · 2023
Later among the works it cites.
Robust test-time adaptation in dynamic scenarios
Longhui Yuan, Binhui Xie, and Shuang Li · 2023
Later among the works it cites.
A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition
Yusheng Dai, Hang Chen, Jun Du, Ruoyu Wang, Shihao Chen, Haotian Wang, and Chin-Hui Lee · 2024
Closest in time.
Exploring missing modality in multimodal egocentric datasets
Merey Ramazanova, Alejandro Pardo, Humam Alwassel, and Bernard Ghanem · 2024
Closest in time.
Test-time adaptation with source based auxiliary tasks
Motasem Alfarra, Alvaro Correia, Bernard Ghanem, and Christos Louizos · 2025
Closest in time.