Fetching the paper…
Reading the bibliography…
Current methods for few-shot action recognition mainly fall into the metric learning framework following ProtoNet, which demonstrates the importance of prototypes.
Visualizing data using t-SNE
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
HMDB: a large video database for human motion recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 2556–2563
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent. In Advances in neural information processing systems (NIPS) , Vol. 29
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Matching networks for one shot learning. In Advances in neural information processing systems (NIPS) , Vol. 29
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition. In European Conference on Computer Vision (ECCV) . 20–36
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 6299–6308
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning (ICML) . 1126–1135
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 5842–5850
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
The effectiveness of data augmentation in image classification using deep learning
Luis Perez and Jason Wang. 2017 · 2017
Earlier work this paper cites.
Learning to compose domain-specific transformations for data augmentation. In Advances in neural information processing systems (NIPS) , Vol. 30
Alexander J Ratner, Henry Ehrenberg, Zeshan Hussain, Jared Dunnmon, and Christopher Ré. 2017 · 2017
Earlier work this paper cites.
Optimization as a model for few-shot learning. In International Conference on Learning Representations (ICLR)
Sachin Ravi and Hugo Larochelle. 2017 · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning. In Advances in neural information processing systems (NIPS) , Vol. 30
Jake Snell, Kevin Swersky, and Richard Zemel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems (NIPS) , Vol. 30
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
How to train your MAML. In International Conference on Learning Representations (ICLR)
Antreas Antoniou, Harrison Edwards, and Amos Storkey. 2018 · 2018
Earlier work this paper cites.
Semantic feature augmentation in few-shot learning
Zitian Chen, Yanwei Fu, Yinda Zhang, Yu-Gang Jiang, Xiangyang Xue, and Leonid Sigal. 2018 · 2018
Earlier work this paper cites.
Few-shot learning with graph neural networks. In International Conference on Learning Representations (ICLR)
Victor Garcia and Joan Bruna. 2018 · 2018
Earlier work this paper cites.
Few-shot human motion prediction via meta-learning. In European Conference on Computer Vision (ECCV) . 432–450
LiangYan Gui, YuXiong Wang, Deva Ramanan, and José MF Moura. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 1199–1208
Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. 2018 · 2018
Cited alongside, same era.
Compound memory networks for few-shot video classification. In European Conference on Computer Vision (ECCV) . 751–766
Linchao Zhu and Yi Yang. 2018 · 2018
Cited alongside, same era.
Tarn: Temporal attentive relation network for few-shot and zero-shot action recognition. In The British Machine Vision Conference (BMVC) . 154
Mina Bishay, Georgios Zoumpourlis, and Ioannis Patras. 2019 · 2019
Cited alongside, same era.
Multi-level semantic feature augmentation for one-shot learning
An ensemble of epoch-wise empirical bayes for few-shot learning. In European Conference on Computer Vision (ECCV) . 404–421
Yaoyao Liu, Bernt Schiele, and Qianru Sun. 2020 · 2020
Later among the works it cites.
Dpgn: Distribution propagation graph network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 13390–13399
Ling Yang, Liangliang Li, Zilun Zhang, Xinyu Zhou, Erjin Zhou, and Yu Liu. 2020 · 2020
Later among the works it cites.
Few-shot learning via embedding adaptation with set-to-set functions. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 8808–8817
Han-Jia Ye, Hexiang Hu, De-Chuan Zhan, and Fei Sha. 2020 · 2020
Later among the works it cites.
Few-shot activity recognition with cross-modal memory network
Lingling Zhang, Xiaojun Chang, Jun Liu, Minnan Luo, Mahesh Prakash, and Alexander G Hauptmann. 2020a · 2020
Later among the works it cites.
What makes multi-modal learning better than single (provably). In Advances in neural information processing systems (NIPS) , Vol. 34. 10944–10956
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zitian Chen, Yanwei Fu, Yinda Zhang, Yu-Gang Jiang, Xiangyang Xue, and Leonid Sigal. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding. In Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, (NAACL-HLT) , Vol. 1. 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Cross attention network for few-shot classification. In Advances in neural information processing systems (NIPS) , Vol. 32
Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. 2019 · 2019
Cited alongside, same era.
Task agnostic meta-learning for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 11719–11727
Muhammad Abdullah Jamal and Guo-Jun Qi. 2019 · 2019
Cited alongside, same era.
Edge-labeling graph neural network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 11–20
Jongmin Kim, Taesup Kim, Sungwoong Kim, and Chang D Yoo. 2019 · 2019
Cited alongside, same era.
Protogan: Towards few shot learning for action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition Workshop (CVPRW)
Sai Kumar Dwivedi, Vikram Gupta, Rahul Mitra, Shuaib Ahmed, and Arjun Jain. 2019 · 2019
Cited alongside, same era.
Meta-learning with differentiable convex optimization. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 10657–10665
Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. 2019 · 2019
Cited alongside, same era.
Finding task-relevant features for few-shot learning by category traversal. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 1–10
Hongyang Li, David Eigen, Samuel Dodge, Matthew Zeiler, and Xiaogang Wang. 2019 · 2019
Cited alongside, same era.
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang. 2021 · 2021
Later among the works it cites.
MASTAF: A Spatio-Temporal Attention Fusion Network for Few-shot Video Classification
Rex Liu, Huanle Zhang, Hamed Pirsiavash, and Xin Liu. 2021 · 2021
Later among the works it cites.
Multimodal prototypical networks for few-shot learning. In IEEE Winter Conference on Applications of Computer Vision (WACV) . 2644–2653
Frederik Pahde, Mihai Puscas, Tassilo Klein, and Moin Nabi. 2021 · 2021
Later among the works it cites.
Temporal-relational crosstransformers for few-shot action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 475–484
Toby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi, and Dima Damen. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning (ICML) . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models. In Advances in neural information processing systems (NIPS) , Vol. 34. 200–212
Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Later among the works it cites.
Actionclip: A new paradigm for video action recognition
Mengmeng Wang, Jiazheng Xing, and Yong Liu. 2021 · 2021
Later among the works it cites.
Few-shot action recognition with prototype-centered attentive learning. In The British Machine Vision Conference (BMVC)
Xiatian Zhu, Antoine Toisoul, Juan-Manuel Perez-Rua, Li Zhang, Brais Martinez, and Tao Xiang. 2021 · 2021
Later among the works it cites.
Compound Prototype Matching for Few-shot Action Recognition. In European Conference on Computer Vision (ECCV)
Yifei Huang, Lijin Yang, and Yoichi Sato. 2022 · 2022
Closest in time.
Denseclip: Language-guided dense prediction with context-aware prompting. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 18082–18091
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. 2022 · 2022
Closest in time.
Spatio-temporal relation modeling for few-shot action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 19958–19967
Anirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, and Bernard Ghanem. 2022 · 2022
Closest in time.
Hybrid Relation Guided Set Matching for Few-shot Action Recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 19948–19957
Xiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang, Zhengrong Zuo, Changxin Gao, Rong Jin, and Nong Sang. 2022 · 2022
Closest in time.
Motion-Modulated Temporal Fragment Alignment Network for Few-Shot Action Recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . 9151–9160
Jiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu, and Yongdong Zhang. 2022 · 2022
Closest in time.
Few-shot action recognition with hierarchical matching and contrastive learning. In European Conference on Computer Vision (ECCV) . 297–313
Sipeng Zheng, Shizhe Chen, and Qin Jin. 2022 · 2022
Closest in time.
CLIP-guided prototype modulating for few-shot action recognition
Xiang Wang, Shiwei Zhang, Jun Cen, Changxin Gao, Yingya Zhang, Deli Zhao, and Nong Sang. 2023a · 2023
Closest in time.
Multimodal Adaptation of CLIP for Few-Shot Action Recognition
Jiazheng Xing, Mengmeng Wang, Xiaojun Hou, Guang Dai, Jingdong Wang, and Yong Liu. 2023a · 2023
Closest in time.