Fetching the paper…
Reading the bibliography…
Zero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data.
Explaining the gibbs sampler
George Casella and Edward I George · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Wordnet:: Similarity-measuring the relatedness of concepts
Ted Pedersen, Siddharth Patwardhan, Jason Michelizzi, et al · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Nltk: the natural language toolkit
Steven Bird · 2006
Earlier work this paper cites.
Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining
Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Guided open vocabulary image captioning with constrained beam search
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Deep compositional captioning: Describing novel object categories without paired training data
Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Senticap: Generating image descriptions with sentiments
Alexander Mathews, Lexing Xie, and Xuming He · 2016
Earlier work this paper cites.
Attend to you: Personalized image captioning with context sequence memory networks
Cesc Chunseong Park, Byeongchang Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Stylenet: Generating attractive visual captions with styles
Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, and Li Deng · 2017
Earlier work this paper cites.
Semantic compositional networks for visual captioning
Zhe Gan, Chuang Gan, Xiaodong He, Yunchen Pu, Kenneth Tran, Jianfeng Gao, Lawrence Carin, and Li Deng · 2017
Earlier work this paper cites.
An empirical study of language cnn for image captioning
Jiuxiang Gu, Gang Wang, Jianfei Cai, and Tsuhan Chen · 2017
Earlier work this paper cites.
High-order attention models for visual question answering
Idan Schwartz, Alexander Schwing, and Tamir Hazan · 2017
Earlier work this paper cites.
Captioning images with diverse objects
Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space
Liwei Wang, Alexander Schwing, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Zero-shot object detection
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Neural baby talk
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2018
Earlier work this paper cites.
Semstyle: Learning to generate stylised image captions using unaligned text
Alexander Mathews, Lexing Xie, and Xuming He · 2018
Earlier work this paper cites.
Diverse beam search for improved description of complex scenes
Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2018
Cited alongside, same era.
Nocaps: Novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson · 2019
Cited alongside, same era.
Sequential latent spaces for modeling the intention during diverse image captioning
Jyoti Aneja, Harsh Agrawal, Dhruv Batra, and Alexander Schwing · 2019
Cited alongside, same era.
Show, control and tell: A framework for generating controllable and grounded captions
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2019
Cited alongside, same era.
Fast, diverse and accurate image captioning guided by part-of-speech
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Later among the works it cites.
Language-driven region pointer advancement for controllable image captioning
Annika Lindh, Robert J Ross, and John D Kelleher · 2020
Later among the works it cites.
X-linear attention networks for image captioning
Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei · 2020
Later among the works it cites.
Show, recall, and tell: Image captioning with recall mechanism
Li Wang, Zechen Bai, Yonghua Zhang, and Hongtao Lu · 2020
Later among the works it cites.
Memcap: Memorizing style knowledge for image captioning
Wentian Zhao, Xinxiao Wu, and Xiaoxun Zhang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander G Schwing, and David Forsyth · 2019
Cited alongside, same era.
Improving text classification with weighted word embeddings via a multi-channel textcnn model
Bao Guo, Chunxia Zhang, Junmin Liu, and Xiaoyi Ma · 2019
Cited alongside, same era.
Mscap: Multi-style image captioning with unpaired stylized text
Longteng Guo, Jing Liu, Peng Yao, Jiangwei Li, and Hanqing Lu · 2019
Cited alongside, same era.
Attention on attention for image captioning
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei · 2019
Cited alongside, same era.
Adaptively aligned image captioning via adaptive attention time
Lun Huang, Wenmin Wang, Yaxian Xia, and Jie Chen · 2019
Cited alongside, same era.
Dense relational captioning: Triple-stream networks for relationship-based captioning
Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, and In So Kweon · 2019
Cited alongside, same era.
Pointing novel objects in image captioning
Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei · 2019
Cited alongside, same era.
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao · 2020
Later among the works it cites.
Human-like controllable image captioning with verb-specific semantic roles
Long Chen, Zhihong Jiang, Jun Xiao, and Wei Liu · 2021
Later among the works it cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Later among the works it cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Later among the works it cites.
Vivo: Visual vocabulary pre-training for novel object captioning
Xiaowei Hu, Xi Yin, Kevin Lin, Lei Zhang, Jianfeng Gao, Lijuan Wang, and Zicheng Liu · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Later among the works it cites.
Multi-scale vision longformer: A new vision transformer for high-resolution image encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
Learning distinct and representative modes for image captioning
Qi Chen, Chaorui Deng, and Qi Wu · 2022
Later among the works it cites.
Injecting semantic concepts into end-to-end image captioning
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang, Zhe Gan, Lijuan Wang, Yezhou Yang, and Zicheng Liu · 2022
Later among the works it cites.
Scaling up vision-language pre-training for image captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Later among the works it cites.
Beyond a pre-trained object detector: Cross-modal textual and visual context for image captioning
Chia-Wen Kuo and Zsolt Kira · 2022
Later among the works it cites.
Grit: Faster and better image captioning transformer using dual visual features
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani · 2022
Later among the works it cites.
A review of generalized zero-shot learning methods
Farhad Pourpanah, Moloud Abdar, Yuxuan Luo, Xinlei Zhou, Ran Wang, Chee Peng Lim, Xi-Zhao Wang, and QM Jonathan Wu · 2022
Later among the works it cites.
Clip models are few-shot learners: Empirical studies on vqa and visual entailment
Haoyu Song, Li Dong, Wei-Nan Zhang, Ting Liu, and Furu Wei · 2022
Later among the works it cites.
Language models can see: Plugging visual controls in text generation
Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier · 2022
Later among the works it cites.
Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic
Yoad Tewel, Yoav Shalev, Idan Schwartz, and Lior Wolf · 2022
Later among the works it cites.
A survey on non-autoregressive generation for neural machine translation and beyond
Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, and Tie-yan Liu · 2022
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Later among the works it cites.