Fetching the paper…
Reading the bibliography…
This work is on training a generative action/video recognition model whose output is a free-form action-specific caption describing the video (rather than an action class label).
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show, adapt and tell: Adversarial training of cross-domain image captioner
Tseng-Hung Chen, Yuan-Hong Liao, Ching-Yao Chuang, Wan-Ting Hsu, Jianlong Fu, and Min Sun · 2017
Earlier work this paper cites.
The Kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Learning a deep embedding model for zero-shot learning
Li Zhang, Tao Xiang, and Shaogang Gong · 2017
Earlier work this paper cites.
A short note about kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Earlier work this paper cites.
Towards universal representation for unseen action recognition
Yi Zhu, Yang Long, Yu Guan, Shawn Newsam, and Ling Shao · 2018
Earlier work this paper cites.
Unsupervised image captioning
Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo · 2019
Earlier work this paper cites.
I know the relationships: Zero-shot action recognition via two-stream graph convolutional networks and knowledge graphs
Junyu Gao, Tianzhu Zhang, and Changsheng Xu · 2019
Earlier work this paper cites.
Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach
Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, and In So Kweon · 2019
Earlier work this paper cites.
Towards unsupervised image captioning with shared multimodal embeddings
Iro Laina, Christian Rupprecht, and Nassir Navab · 2019
Earlier work this paper cites.
Quality estimation for image captions based on large-scale human evaluations
Tomer Levinboim, Ashish V Thapliyal, Piyush Sharma, and Radu Soricut · 2019
Earlier work this paper cites.
TSM: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Pseudo-labeling and confirmation bias in deep semi-supervised learning
Eric Arazo, Diego Ortego, Paul Albert, Noel E O’Connor, and Kevin McGuinness · 2020
Earlier work this paper cites.
Rethinking zero-shot video classification: End-to-end training for realistic applications
Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, and Krzysztof Chalupka · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
All about knowledge graphs for actions
Pallabi Ghosh, Nirat Saini, Larry S Davis, and Abhinav Shrivastava · 2020
Cited alongside, same era.
Spatio-temporal graph for video captioning with knowledge distillation
Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Closest in time.
Fitclip: Refining large-scale pretrained image-text models for zero-shot video understanding tasks
Santiago Castro and Fabian Caba Heilbron · 2022
Closest in time.
Zero-shot action recognition with transformer-based video semantic embedding
Keval Doshi and Yasin Yilmaz · 2022
Closest in time.
Learning to prompt for open-vocabulary object detection with vision-language model
Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li · 2022
Closest in time.
Global semantic descriptors for zero-shot action recognition
Valter Estevam, Rayson Laroca, Helio Pedrini, and David Menotti · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zero-shot action recognition from diverse object-scene compositions
Carlo Bretti and Pascal Mettes · 2021
Cited alongside, same era.
Elaborative rehearsal for zero-shot action recognition
Shizhe Chen and Dong Huang · 2021
Cited alongside, same era.
Self-distillation for few-shot image captioning
Xianyu Chen, Ming Jiang, and Qi Zhao · 2021
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano · 2021
Cited alongside, same era.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Cited alongside, same era.
Closest in time.
Bridging video-text retrieval with multiple choice questions
Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li, Ying Shan, Xiaohu Qie, and Ping Luo · 2022
Closest in time.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Closest in time.
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Closest in time.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Closest in time.
Disentangled action recognition with knowledge bases
Zhekun Luo, Shalini Ghosh, Devin Guillory, Keizo Kato, Trevor Darrell, and Huijuan Xu · 2022
Closest in time.
Universal prototype transport for zero-shot action recognition and localization
Pascal Mettes · 2022
Closest in time.
Simple open-vocabulary object detection with vision transformers
Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, et al · 2022
Closest in time.
Expanding language-image pretrained models for general video recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling · 2022
Closest in time.
Alignment-uniformity aware representation learning for zero-shot video classification
Shi Pu, Kaili Zhao, and Mao Zheng · 2022
Closest in time.
Multimodal open-vocabulary video classification via pre-trained vision and language models
Rui Qian, Yeqing Li, Zheng Xu, Ming-Hsuan Yang, Serge Belongie, and Yin Cui · 2022
Closest in time.
GIT: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Closest in time.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Closest in time.