Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Learning to recognize daily actions using gaze
Alireza Fathi, Yin Li, James M Rehg, et al · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
First-person activity recognition: What are they doing to me?
Michael S Ryoo and Larry Matthies · 2013
Earlier work this paper cites.
You-do, i-learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video
Dima Damen, Teesid Leelasawassuk, Osian Haines, Andrew Calway, and Walterio W Mayol-Cuevas · 2014
Earlier work this paper cites.
Action and interaction recognition in first-person videos
Sanath Narayan, Mohan S Kankanhalli, and Kalpathi R Ramakrishnan · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Discovering task relevant objects and their modes of interaction from multi-user egocentric video
A Calway, W Mayol-Cuevas, D Damen, O Haines, and T Leelasawassuk · 2015
Earlier work this paper cites.
Predicting important objects for egocentric video summarization
Yong Jae Lee and Kristen Grauman · 2015
Earlier work this paper cites.
Weakly-shared deep transfer networks for heterogeneous-domain knowledge propagation
Xiangbo Shu, Guo-Jun Qi, Jinhui Tang, and Jingdong Wang · 2015
Earlier work this paper cites.
Personalized age progression with aging dictionary
Xiangbo Shu, Jinhui Tang, Hanjiang Lai, Luoqi Liu, and Shuicheng Yan · 2015
Earlier work this paper cites.
Summarization of egocentric videos: A comprehensive survey
Ana Garcia Del Molino, Cheston Tan, Joo-Hwee Lim, and Ah-Hwee Tan · 2016
Earlier work this paper cites.
Multimodal multi-stream deep learning for egocentric activity recognition
Sibo Song, Vijay Chandrasekhar, Bappaditya Mandal, Liyuan Li, Joo-Hwee Lim, Giduthuri Sateesh Babu, Phyo Phyo San, and Ngai-Man Cheung · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Spatiotemporal pyramid network for video action recognition
Yunbo Wang, Mingsheng Long, Jianmin Wang, and Philip S Yu · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Earlier work this paper cites.
Leveraging uncertainty to rethink loss functions and evaluation measures for egocentric action anticipation
Antonino Furnari, Sebastiano Battiato, and Giovanni Maria Farinella · 2018
Earlier work this paper cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Earlier work this paper cites.
Actor and observer: Joint modeling of first and third-person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Attention is all we need: Nailing down object-centric attention for egocentric activity recognition
Swathikiran Sudhakaran and Oswald Lanz · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Dynamic temporal pyramid network: A closer look at multi-scale modeling for activity detection
Da Zhang, Xiyang Dai, and Yuan-Fang Wang · 2018
Earlier work this paper cites.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford Alec, Wu Jeffrey, Child Rewon, Luan David, Amodei Dario, and et al · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Multitask learning to improve egocentric action recognition
Georgios Kapidis, Ronald Poppe, Elsbeth van Dam, Lucas Noldus, and Remco Veltkamp · 2019
Earlier work this paper cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Earlier work this paper cites.
Learning spatiotemporal attention for egocentric action recognition
Minlong Lu, Danping Liao, and Ze-Nian Li · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Hierarchical long short-term concurrent memory for human interaction recognition
Xiangbo Shu, Jinhui Tang, Guo-Jun Qi, Wei Liu, and Jian Yang · 2019
Earlier work this paper cites.
Lsta: Long short-term attention for egocentric action recognition
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Earlier work this paper cites.
Coherence constrained graph lstm for group activity recognition
Jinhui Tang, Xiangbo Shu, Rui Yan, and Liyan Zhang · 2019
Earlier work this paper cites.
Stacked memory network for video summarization
Junbo Wang, Wei Wang, Zhiyong Wang, Liang Wang, Dagan Feng, and Tieniu Tan · 2019
Earlier work this paper cites.
Hallucinating idt descriptors and i3d optical flow features for action recognition with cnns
Lei Wang, Piotr Koniusz, and Du Q Huynh · 2019
Earlier work this paper cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Earlier work this paper cites.
Progressive instance-aware feature learning for compositional action recognition
Rui Yan, Lingxi Xie, Xiangbo Shu, Liyan Zhang, and Jinhui Tang · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, and et al · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Rolling-unrolling lstms for action anticipation from first-person video
Antonino Furnari and Giovanni Maria Farinella · 2020
Earlier work this paper cites.
Mutual context network for jointly estimating egocentric gaze and action
Yifei Huang, Minjie Cai, Zhenqiang Li, Feng Lu, and Yoichi Sato · 2020
Earlier work this paper cites.
Host–parasite: Graph lstm-in-lstm for group activity recognition
Xiangbo Shu, Liyan Zhang, Yunlian Sun, and Jinhui Tang · 2020
Cited alongside, same era.
Learning video representations from textual web supervision
Jonathan C Stroud, Zhichao Lu, Chen Sun, Jia Deng, Rahul Sukthankar, Cordelia Schmid, and David A Ross · 2020
Cited alongside, same era.
Symbiotic attention with privileged information for egocentric action recognition
Xiaohan Wang, Yu Wu, Linchao Zhu, and Yi Yang · 2020
Cited alongside, same era.
Symbiotic attention for egocentric action recognition with object-centric alignment
Xiaohan Wang, Linchao Zhu, Yu Wu, and Yi Yang · 2020
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Expansion-squeeze-excitation fusion network for elderly activity recognition
Xiangbo Shu, Jiawen Yang, Rui Yan, and Yan Song · 2022
Later among the works it cites.
Distance matters in human-object interaction detection
Guangzhi Wang, Yangyang Guo, Yongkang Wong, and Mohan Kankanhalli · 2022
Later among the works it cites.
Continuous multi-view human action recognition
Qiang Wang, Gan Sun, Jiahua Dong, Qianqian Wang, and Zhengming Ding · 2022
Later among the works it cites.
Internvideo: General video foundation models via generative and discriminative learning
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al · 2022
Later among the works it cites.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Cited alongside, same era.
Space-time mixing attention for video transformer
Adrian Bulat, Juan Manuel Perez Rua, Swathikiran Sudhakaran, Brais Martinez, and Georgios Tzimiropoulos · 2021
Cited alongside, same era.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2021
Cited alongside, same era.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Cited alongside, same era.
You only look at one sequence: Rethinking transformer in vision through object detection
Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu · 2021
Cited alongside, same era.
Multimodal global relation knowledge distillation for egocentric action anticipation
Yi Huang, Xiaoshan Yang, and Changsheng Xu · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Look less think more: Rethinking compositional action recognition
Rui Yan, Peng Huang, Xiangbo Shu, Junhao Zhang, Yonghua Pan, and Jinhui Tang · 2022
Later among the works it cites.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Is an object-centric video representation beneficial for transfer?
Chuhan Zhang, Ankush Gupta, and Andrew Zisserman · 2022
Later among the works it cites.
Hierarchical few-shot object detection: Problem, benchmark and method
Lu Zhang, Yang Wang, Jiaogen Zhou, Chenbo Zhang, Yinglu Zhang, Jihong Guan, Yatao Bian, and Shuigeng Zhou · 2022
Later among the works it cites.
Hiervl: Learning hierarchical video-language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman · 2023
Later among the works it cites.
Hgformer: Hierarchical grouping transformer for domain generalized semantic segmentation
Jian Ding, Nan Xue, Gui-Song Xia, Bernt Schiele, and Dengxin Dai · 2023
Later among the works it cites.
Gpt4image: Can large pre-trained models help vision models on perception tasks?
Ning Ding, Yehui Tang, Zhongqian Fu, Chao Xu, Kai Han, and Yunhe Wang · 2023
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2023
Later among the works it cites.
Egotv: Egocentric task verification from natural language task descriptions
Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, and Ruta Desai · 2023
Later among the works it cites.
Egohumans: An egocentric 3d multi-human benchmark
Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani · 2023
Later among the works it cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Videochat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao · 2023
Later among the works it cites.
A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging
Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lingqiao Liu, Lei Wang, Leyang Cui, Zhaopeng Tu, Longyue Wang, and Luping Zhou · 2023
Later among the works it cites.
Mm-vid: Advancing video understanding with gpt-4v (ision)
Kevin Lin, Faisal Ahmed, Linjie Li, Chung-Ching Lin, Ehsan Azarnasab, Zhengyuan Yang, Jianfeng Wang, Lin Liang, Zicheng Liu, Yumao Lu, et al · 2023
Later among the works it cites.
Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou · 2023
Later among the works it cites.
Visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Attention-driven appearance-motion fusion network for action recognition
Shaocan Liu and Xin Ma · 2023
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Gpt-4v (ision) as a social media analysis engine
Hanjia Lyu, Jinfa Huang, Daoan Zhang, Yongsheng Yu, Xinyi Mou, Jinsheng Pan, Zhengyuan Yang, Zhongyu Wei, and Jiebo Luo · 2023
Later among the works it cites.
Pip-net: Patch-based intuitive prototypes for interpretable image classification
Meike Nauta, Jörg Schlötterer, Maurice van Keulen, and Christin Seifert · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
Skeleton-based action recognition through contrasting two-stream spatial-temporal networks
Chen Pang, Xuequan Lu, and Lei Lyu · 2023
Later among the works it cites.
Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang · 2023
Later among the works it cites.
Energy-based temporal summarized attentive network for zero-shot action recognition
Cheng Qi, Zhiyong Feng, Meng Xing, Yong Su, Jinqing Zheng, and Yiming Zhang · 2023
Later among the works it cites.
Chatgpt-powered hierarchical comparisons for image classification
Zhiyuan Ren, Yiyang Su, and Xiaoming Liu · 2023
Later among the works it cites.
Exploring ocr capabilities of gpt-4v (ision): A quantitative and in-depth evaluation
Yongxin Shi, Dezhi Peng, Wenhui Liao, Zening Lin, Xinhong Chen, Chongyu Liu, Yuyi Zhang, and Lianwen Jin · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Dear-net: Learning diversities for skeleton-based early action recognition
Rui Wang, Jun Liu, Qiuhong Ke, Duo Peng, and Yinjie Lei · 2023
Later among the works it cites.
Detecting everything in the open world: Towards universal object detection
Zhenyu Wang, Yali Li, Xi Chen, Ser-Nam Lim, Antonio Torralba, Hengshuang Zhao, and Shengjin Wang · 2023
Later among the works it cites.
On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving
Licheng Wen, Xuemeng Yang, Daocheng Fu, Xiaofeng Wang, Pinlong Cai, Xin Li, Tao Ma, Yingxuan Li, Linran Xu, Dengke Shang, et al · 2023
Later among the works it cites.
Cap4video: What can auxiliary captions do for text-video retrieval?
Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang · 2023
Later among the works it cites.
Revisiting classifier: Transferring vision-language models for video recognition
Wenhao Wu, Zhun Sun, and Wanli Ouyang · 2023
Later among the works it cites.
Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models
Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, and Wanli Ouyang · 2023
Later among the works it cites.
Gpt4vis: What can gpt-4 do for zero-shot visual recognition?
Wenhao Wu, Huanjin Yao, Mengxi Zhang, Yuxin Song, Wanli Ouyang, and Jingdong Wang · 2023
Later among the works it cites.
Lingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang, and Jian Yang · 2023
Later among the works it cites.
Aim: Adapting image models for efficient video action recognition
Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li · 2023
Later among the works it cites.
Performance of multimodal gpt-4v on usmle with image: Potential for imaging diagnostic support with explanations
Zhichao Yang, Zonghai Yao, Mahbuba Tasmin, Parth Vashisht, Won Seok Jang, Beining Wang, Dan Berlowitz, and Hong Yu · 2023
Later among the works it cites.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Later among the works it cites.
Exploring recommendation capabilities of gpt-4v (ision): A preliminary case study
Peilin Zhou, Meng Cao, You-Liang Huang, Qichen Ye, Peiyan Zhang, Junling Liu, Yueqi Xie, Yining Hua, and Jaeboum Kim · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Transferring vision-language models for visual recognition: A classifier perspective
Wenhao Wu, Zhun Sun, Yuxin Song, Jingdong Wang, and Wanli Ouyang · 2024
Closest in time.