Fetching the paper…
Reading the bibliography…
Connecting Vision and Language plays an essential role in Generative Intelligence.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , 1992
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in ACL , 2002
2002
Earlier work this paper cites.
J.-Y. Pan, H.-J. Yang, P. Duygulu, and C. Faloutsos, “Automatic image captioning,” in ICME , 2004
2004
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in ACL Workshops , 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in ACL Workshops , 2005
2005
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in ICVGIP , 2008
2008
Earlier work this paper cites.
P. Koehn, Statistical Machine Translation . Cambridge University Press, 2009
2009
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in ECCV , 2010
2010
Earlier work this paper cites.
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu, “I2T: Image parsing to text description,” Proceedings of the IEEE , 2010
2010
Earlier work this paper cites.
A. Aker and R. Gaizauskas, “Generating image descriptions using dependency relational patterns,” in ACL , 2010
2010
Earlier work this paper cites.
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona, “Caltech-UCSD Birds 200,” California Institute of Technology, Tech. Rep., 2010
2010
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” in NeurIPS , 2011
2011
Earlier work this paper cites.
Y. Yang, C. Teo, H. Daumé III, and Y. Aloimonos, “Corpus-guided sentence generation of natural images,” in EMNLP , 2011
2011
Earlier work this paper cites.
S. Li, G. Kulkarni, T. Berg, A. Berg, and Y. Choi, “Composing simple image descriptions using web-scale n-grams,” in CoNLL , 2011
2011
Earlier work this paper cites.
A. Gupta, Y. Verma, and C. Jawahar, “Choosing linguistics over vision to describe images,” in AAAI , 2012
2012
Earlier work this paper cites.
M. Mitchell, J. Dodge, A. Goyal, K. Yamaguchi, K. Stratos, X. Han, A. Mensch, A. Berg, T. Berg, and H. Daumé III, “Midge: Generating image descriptions from computer vision detections,” in ACL , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NeurIPS , 2012
2012
Earlier work this paper cites.
Y. Feng and M. Lapata, “Automatic Caption Generation for News Images,” IEEE Trans. PAMI , vol. 35, no. 4, pp. 797–812, 2012
2012
Earlier work this paper cites.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, “DeViSE: a deep visual-semantic embedding model,” in NeurIPS , 2013
2013
Earlier work this paper cites.
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “BabyTalk: Understanding and generating simple image descriptions,” IEEE Trans. PAMI , 2013
2013
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” JAIR , 2013
2013
Earlier work this paper cites.
R. Kiros, R. Salakhutdinov, and R. S. Zemel, “Unifying visual-semantic embeddings with multimodal neural language models,” in NeurIPS Workshops , 2014
2014
Earlier work this paper cites.
A. Karpathy, A. Joulin, and L. Fei-Fei, “Deep fragment embeddings for bidirectional image sentence mapping,” in NeurIPS , 2014
2014
Earlier work this paper cites.
P. Kuznetsova, V. Ordonez, T. L. Berg, and Y. Choi, “Treetalk: Composition and compression of trees for image descriptions,” TACL , vol. 2, pp. 351–362, 2014
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR , 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common Objects in Context,” in ECCV , 2014
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL , 2014
2014
Earlier work this paper cites.
A. Ardila, B. Bernal, and M. Rosselli, “Language and visual perception associations: meta-analytic connectivity modeling of Brodmann Area 37,” Behavioural Neurology , 2015
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in CVPR , 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR , 2015
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR , 2015
2015
Earlier work this paper cites.
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille, “Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN),” in ICLR , 2015
2015
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
X. Chen and C. Lawrence Zitnick, “Mind’s Eye: A Recurrent Visual Representation for Image Caption Generation,” in CVPR , 2015
2015
Earlier work this paper cites.
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt et al. , “From captions to visual concepts and back,” in CVPR , 2015
2015
Earlier work this paper cites.
X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars, “Guiding the Long-Short Term Memory model for Image Caption Generation,” in ICCV , 2015
2015
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based Image Description Evaluation,” in CVPR , 2015
2015
Earlier work this paper cites.
M. Kusner, Y. Sun, N. Kolkin, and K. Weinberger, “From word embeddings to document distances,” in ICML , 2015
2015
Earlier work this paper cites.
D. Elliott, S. Frank, and E. Hasler, “Multilingual image description with neural sequence models,” ICLR , 2015
2015
Earlier work this paper cites.
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank, “Automatic description generation from images: A survey of models, datasets, and evaluation measures,” JAIR , vol. 55, pp. 409–442, 2016
2016
Earlier work this paper cites.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in CVPR , 2016
2016
Earlier work this paper cites.
Q. Wu, C. Shen, L. Liu, A. Dick, and A. Van Den Hengel, “What Value Do Explicit High Level Concepts Have in Vision to Language Problems?” in CVPR , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
Z. Yang, Y. Yuan, Y. Wu, W. W. Cohen, and R. R. Salakhutdinov, “Review Networks for Caption Generation,” in NeurIPS , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba, “Sequence level training with recurrent neural networks,” in ICLR , 2016
2016
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic Propositional Image Caption Evaluation,” in ECCV , 2016
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural Machine Translation of Rare Words with Subword Units,” in ACL , 2016
2016
Earlier work this paper cites.
S. Reed, Z. Akata, H. Lee, and B. Schiele, “Learning deep representations of fine-grained visual descriptions,” in CVPR , 2016
2016
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “YFCC100M: The new data in multimedia research,” Communications of the ACM , vol. 59, no. 2, pp. 64–73, 2016
2016
Earlier work this paper cites.
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell, “Generating visual explanations,” in ECCV , 2016
2016
Earlier work this paper cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, “Generative adversarial text to image synthesis,” in ICML , 2016
2016
Earlier work this paper cites.
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell, “Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data,” in CVPR , 2016
2016
Earlier work this paper cites.
J. Johnson, A. Karpathy, and L. Fei-Fei, “DenseCap: Fully convolutional Localization Networks for Dense Captioning,” in CVPR , 2016
2016
Earlier work this paper cites.
T. Miyazaki and N. Shimizu, “Cross-lingual image caption generation,” in ACM Multimedia , 2016
2016
Earlier work this paper cites.
D. Elliott, S. Frank, K. Sima’an, and L. Specia, “Multi30K: Multilingual English-German Image Descriptions,” in ACL Workshops , 2016
2016
Earlier work this paper cites.
M. Hodosh and J. Hockenmaier, “Focused evaluation for image description with binary forced-choice tasks,” in ACL Workshops , 2016
2016
Earlier work this paper cites.
J. Gu, G. Wang, J. Cai, and T. Chen, “An Empirical Study of Language CNN for Image Captioning,” in ICCV , 2017
2017
Earlier work this paper cites.
F. Chen, R. Ji, J. Su, Y. Wu, and Y. Wu, “StructCap: Structured Semantic Embedding for Image Captioning,” in ACM Multimedia , 2017
2017
Earlier work this paper cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in CVPR , 2017
2017
Earlier work this paper cites.
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei, “Boosting image captioning with attributes,” in ICCV , 2017
2017
Earlier work this paper cites.
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng, “Semantic Compositional Networks for Visual Captioning,” in CVPR , 2017
2017
Earlier work this paper cites.
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in CVPR , 2017
2017
Earlier work this paper cites.
Y. Wang, Z. Lin, X. Shen, S. Cohen, and G. W. Cottrell, “Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition,” in CVPR , 2017
2017
Earlier work this paper cites.
L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T.-S. Chua, “SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning,” in CVPR , 2017
2017
Earlier work this paper cites.
H. R. Tavakoli, R. Shetty, A. Borji, and J. Laaksonen, “Paying attention to descriptions generated by image captioning models,” in ICCV , 2017
2017
Earlier work this paper cites.
V. Ramanishka, A. Das, J. Zhang, and K. Saenko, “Top-down visual saliency guided by captions,” in CVPR , 2017
2017
Earlier work this paper cites.
——, “Faster R-CNN: towards real-time object detection with region proposal networks,” IEEE Trans. PAMI , vol. 39, no. 6, pp. 1137–1149, 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei, “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations,” IJCV , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
M. Pedersoli, T. Lucas, C. Schmid, and J. Verbeek, “Areas of Attention for Image Captioning,” in ICCV , 2017
2017
Earlier work this paper cites.
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in ICLR , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
Z. Ren, X. Wang, N. Zhang, X. Lv, and L.-J. Li, “Deep reinforcement learning-based image captioning with embedding reward,” in CVPR , 2017
2017
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved Image Captioning via Policy Gradient Optimization of SPIDEr,” in ICCV , 2017
2017
Earlier work this paper cites.
L. Zhang, F. Sung, F. Liu, T. Xiang, S. Gong, Y. Yang, and T. M. Hospedales, “Actor-Critic Sequence Training for Image Captioning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
A. Ramisa, F. Yan, F. Moreno-Noguer, and K. Mikolajczyk, “BreakingNews: Article Annotation by Image and Text Processing,” IEEE Trans. PAMI , vol. 40, no. 5, pp. 1072–1085, 2017
2017
Earlier work this paper cites.
T.-H. Chen, Y.-H. Liao, C.-Y. Chuang, W.-T. Hsu, J. Fu, and M. Sun, “Show, adapt and tell: Adversarial training of cross-domain image captioner,” in ICCV , 2017
2017
Earlier work this paper cites.
R. Shetty, M. Rohrbach, L. Anne Hendricks, M. Fritz, and B. Schiele, “Speaking the same language: Matching machine to human captions by adversarial training,” in ICCV , 2017
2017
Earlier work this paper cites.
M. Kilickaya, A. Erdem, N. Ikizler-Cinbis, and E. Erdem, “Re-evaluating automatic metrics for image captioning,” in ACL , 2017
2017
Earlier work this paper cites.
B. Dai, S. Fidler, R. Urtasun, and D. Lin, “Towards Diverse and Natural Image Descriptions via a Conditional GAN,” in ICCV , 2017
2017
Earlier work this paper cites.
S. Venugopalan, L. A. Hendricks, M. Rohrbach, R. Mooney, T. Darrell, and K. Saenko, “Captioning Images with Diverse Objects,” in CVPR , 2017
2017
Earlier work this paper cites.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects,” in CVPR , 2017
2017
Cited alongside, same era.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Guided Open Vocabulary Image Captioning with Constrained Beam Search,” in EMNLP , 2017
2017
Cited alongside, same era.
L. Yang, K. Tang, J. Yang, and L.-J. Li, “Dense captioning with joint inference and visual context,” in CVPR , 2017
2017
Cited alongside, same era.
J. Krause, J. Johnson, R. Krishna, and L. Fei-Fei, “A hierarchical approach for generating descriptive image paragraphs,” in CVPR , 2017
2017
Cited alongside, same era.
X. Liang, Z. Hu, H. Zhang, C. Gan, and E. P. Xing, “Recurrent topic-transition GAN for visual paragraph generation,” in ICCV , 2017
2017
Cited alongside, same era.
K. Shuster, S. Humeau, H. Hu, A. Bordes, and J. Weston, “Engaging image captioning via personality,” in CVPR , 2019
2019
Later among the works it cites.
Y. Zheng, Y. Li, and S. Wang, “Intention oriented image captions with guiding objects,” in CVPR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Sharif, U. Nadeem, S. A. A. Shah, M. Bennamoun, and W. Liu, “Vision to Language: Methods, Metrics and Datasets,” in Machine Learning Paradigms , 2020, pp. 9–62
2020
Later among the works it cites.
H. Sharma, M. Agrahari, S. K. Singh, M. Firoj, and R. K. Mishra, “Image captioning: a comprehensive survey,” in PARC , 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Dai and D. Lin, “Contrastive learning for image captioning,” in NeurIPS , 2017
2017
Cited alongside, same era.
L. Wang, A. G. Schwing, and S. Lazebnik, “Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space,” in NeurIPS , 2017
2017
Cited alongside, same era.
W. Lan, X. Li, and J. Dong, “Fluency-guided cross-lingual image captioning,” in ACM Multimedia , 2017
2017
Cited alongside, same era.
S. Wu, J. Wieland, O. Farivar, and J. Schiller, “Automatic alt-text: Computer-generated image descriptions for blind users on a social network service,” in CSCW , 2017
2017
Cited alongside, same era.
C. Chunseong Park, B. Kim, and G. Kim, “Attend to you: Personalized image captioning with context sequence memory networks,” in CVPR , 2017
2017
Cited alongside, same era.
C. Gan, Z. Gan, X. He, J. Gao, and L. Deng, “StyleNet: Generating Attractive Visual Captions with Styles,” in CVPR , 2017
2017
Cited alongside, same era.
S. Bai and S. An, “A survey on automatic image caption generation,” Neurocomputing , vol. 311, pp. 291–304, 2018
2018
Cited alongside, same era.
Later among the works it cites.
L. Wang, Z. Bai, Y. Zhang, and H. Lu, “Show, Recall, and Tell: Image Captioning with Recall Mechanism,” in AAAI , 2020
2020
Later among the works it cites.
Z. Shi, X. Zhou, X. Qiu, and X. Zhu, “Improving Image Captioning with Better Use of Captions,” in ACL , 2020
2020
Later among the works it cites.
L. Guo, J. Liu, X. Zhu, P. Yao, S. Lu, and H. Lu, “Normalized and Geometry-Aware Self-Attention Network for Image Captioning,” in CVPR , 2020
2020
Later among the works it cites.
Y. Pan, T. Yao, Y. Li, and T. Mei, “X-Linear Attention Networks for Image Captioning,” in CVPR , 2020
2020
Later among the works it cites.
M. Cornia, M. Stefanini, L. Baraldi, and R. Cucchiara, “Meshed-Memory Transformer for Image Captioning,” in CVPR , 2020
2020
Later among the works it cites.
S. He, W. Liao, H. R. Tavakoli, M. Yang, B. Rosenhahn, and N. Pugeault, “Image captioning through image transformer,” in ACCV , 2020
2020
Later among the works it cites.
F. Liu, X. Ren, X. Wu, S. Ge, W. Fan, Y. Zou, and X. Sun, “Prophet Attention: Predicting Attention with Future Attention,” in NeurIPS , 2020
2020
Later among the works it cites.
M. Cornia, L. Baraldi, and R. Cucchiara, “SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability,” in ICRA , 2020
2020
Later among the works it cites.
H. Jiang, I. Misra, M. Rohrbach, E. Learned-Miller, and X. Chen, “In defense of grid features for visual question answering,” in CVPR , 2020
2020
Later among the works it cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in ECCV , 2020
2020
Later among the works it cites.
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. J. Corso, and J. Gao, “Unified Vision-Language Pre-Training for Image Captioning and VQA,” in AAAI , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
L. Guo, J. Liu, X. Zhu, X. He, J. Jiang, and H. Lu, “Non-autoregressive image captioning with counterfactuals-critical multi-agent learning,” IJCAI , 2020
2020
Later among the works it cites.
Z. Fei, “Iterative Back Modification for Faster Image Captioning,” in ACM Multimedia , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
D. Gurari, Y. Zhao, M. Zhang, and N. Bhattacharya, “Captioning Images Taken by People Who Are Blind,” in ECCV , 2020
2020
Later among the works it cites.
X. Yang, H. Zhang, D. Jin, Y. Liu, C.-H. Wu, J. Tan, D. Xie, J. Wang, and X. Wang, “Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards,” in ECCV , 2020
2020
Later among the works it cites.
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “TextCaps: a Dataset for Image Captioning with Reading Comprehension,” in ECCV , 2020
2020
Later among the works it cites.
J. Pont-Tuset, J. Uijlings, S. Changpinyo, R. Soricut, and V. Ferrari, “Connecting vision and language with localized narratives,” in ECCV , 2020
2020
Later among the works it cites.
O. Caglayan, P. Madhyastha, and L. Specia, “Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale,” in COLING , 2020
2020
Later among the works it cites.
Q. Wang, J. Wan, and A. B. Chan, “On Diversity in Image Captioning: Metrics and Methods,” IEEE Trans. PAMI , 2020
2020
Later among the works it cites.
Z. Wang, B. Feng, K. Narasimhan, and O. Russakovsky, “Towards unique and informative captioning of images,” in ECCV , 2020
2020
Later among the works it cites.
R. Bigazzi, F. Landi, M. Cornia, S. Cascianelli, L. Baraldi, and R. Cucchiara, “Explore and Explain: Self-supervised Navigation and Recounting,” in ICPR , 2020
2020
Later among the works it cites.
H. Lee, S. Yoon, F. Dernoncourt, D. S. Kim, T. Bui, and K. Jung, “ViLBERTScore: Evaluating Image Caption Using Vision-and-Language BERT,” in EMNLP Workshops , 2020
2020
Later among the works it cites.
Y. Yi, H. Deng, and J. Hu, “Improving image captioning evaluation by considering inter references variance,” in ACL , 2020
2020
Later among the works it cites.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “BERTScore: Evaluating Text Generation with BERT,” in ICLR , 2020
2020
Later among the works it cites.
X. Hu, X. Yin, K. Lin, L. Wang, L. Zhang, J. Gao, and Z. Liu, “VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning,” AAAI , 2020
2020
Later among the works it cites.
D. Guo, Y. Wang, P. Song, and M. Wang, “Recurrent relational memory network for unsupervised image captioning,” IJCAI , 2020
2020
Later among the works it cites.
R. Del Chiaro, B. Twardowski, A. D. Bagdanov, and J. van de Weijer, “RATT: Recurrent Attention to Transient Tasks for Continual Image Captioning,” in NeurIPS , 2020
2020
Later among the works it cites.
J. Wang, J. Tang, and J. Luo, “Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image Captioning,” in ACM Multimedia , 2020
2020
Later among the works it cites.
X. Shi, X. Yang, J. Gu, S. Joty, and J. Cai, “Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning,” in ECCV , 2020
2020
Later among the works it cites.
S. Mahajan and S. Roth, “Diverse image captioning with context-object split latent spaces,” in NeurIPS , 2020
2020
Later among the works it cites.
A. Tran, A. Mathews, and L. Xie, “Transform and Tell: Entity-Aware News Image Captioning,” in CVPR , 2020
2020
Later among the works it cites.
W. Zhang, Y. Ying, P. Lu, and H. Zha, “Learning Long-and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption,” in AAAI , 2020
2020
Later among the works it cites.
W. Zhao, X. Wu, and X. Zhang, “MemCap: Memorizing style knowledge for image captioning,” in AAAI , 2020
2020
Later among the works it cites.
S. Chen, Q. Jin, P. Wang, and Q. Wu, “Say as you wish: Fine-grained control of image caption generation with abstract scene graphs,” in CVPR , 2020
2020
Later among the works it cites.
Y. Zhong, L. Wang, J. Chen, D. Yu, and Y. Li, “Comprehensive image captioning via scene graph decomposition,” in ECCV , 2020
2020
Later among the works it cites.
C. Deng, N. Ding, M. Tan, and Q. Wu, “Length-controllable image captioning,” in ECCV , 2020
2020
Later among the works it cites.
F. Sammani and L. Melas-Kyriazi, “Show, edit and tell: A framework for editing image captions,” in CVPR , 2020
2020
Later among the works it cites.
M. Alikhani, P. Sharma, S. Li, R. Soricut, and M. Stone, “Cross-modal Coherence Modeling for Caption Generation,” in ACL , 2020
2020
Later among the works it cites.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “UNITER: UNiversal Image-TExt Representation Learning,” in ECCV , 2020
2020
Later among the works it cites.
J. Ji, Y. Luo, X. Sun, F. Chen, G. Luo, Y. Wu, Y. Gao, and R. Ji, “Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network,” in AAAI , 2021
2021
Closest in time.
Y. Luo, J. Ji, X. Sun, L. Cao, Y. Wu, F. Huang, C.-W. Lin, and R. Ji, “Dual-Level Collaborative Transformer for Image Captioning,” in AAAI , 2021
2021
Closest in time.
X. Zhang, X. Sun, Y. Luo, J. Ji, Y. Zhou, Y. Wu, F. Huang, and R. Ji, “RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words,” in CVPR , 2021
2021
Closest in time.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” ICLR , 2021
2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in ICML , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “VinVL: Revisiting visual representations in vision-language models,” in CVPR , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
K. Desai and J. Johnson, “VirTex: Learning Visual Representations From Textual Annotations,” in CVPR , 2021
2021
Closest in time.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts,” in CVPR , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in ICML , 2021
2021
Closest in time.
2021
Closest in time.
S. Wang, Z. Yao, R. Wang, Z. Wu, and X. Chen, “FAIEr: Fidelity and Adequacy Ensured Image Caption Evaluation,” in CVPR , 2021
2021
Closest in time.
H. Lee, S. Yoon, F. Dernoncourt, T. Bui, and K. Jung, “UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning,” in ACL , 2021
2021
Closest in time.
I. J. Unanue, J. Parnell, and M. Piccardi, “BERTTune: Fine-Tuning Neural Machine Translation with BERTScore,” ACL , 2021
2021
Closest in time.
2021
Closest in time.
H. Ben, Y. Pan, Y. Li, T. Yao, R. Hong, M. Wang, and T. Mei, “Unpaired Image Captioning with Semantic-Constrained Self-Learning,” IEEE Trans. Multimedia , 2021
2021
Closest in time.
Z. Yang, Y. Lu, J. Wang, X. Yin, D. Florencio, L. Wang, C. Zhang, L. Zhang, and J. Luo, “TAP: Text-Aware Pre-training for Text-VQA and Text-Caption,” in CVPR , 2021
2021
Closest in time.
J. Wang, J. Tang, M. Yang, X. Bai, and J. Luo, “Improving OCR-based Image Captioning by Incorporating Geometrical Relationship,” in CVPR , 2021
2021
Closest in time.
Q. Zhu, C. Gao, P. Wang, and Q. Wu, “Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps,” in AAAI , 2021
2021
Closest in time.
G. Xu, S. Niu, M. Tan, Y. Luo, Q. Du, and Q. Wu, “Towards Accurate Text-based Image Captioning with Content Diversity Exploration,” in CVPR , 2021
2021
Closest in time.
Q. Huang, Y. Liang, J. Wei, C. Yi, H. Liang, H.-f. Leung, and Q. Li, “Image Difference Captioning with Instance-Level Fine-Grained Feature Representation,” IEEE Trans. Multimedia , 2021
2021
Closest in time.
H. Kim, J. Kim, H. Lee, H. Park, and G. Kim, “Viewpoint-Agnostic Change Captioning With Cycle Consistency,” in ICCV , 2021
2021
Closest in time.
M. Hosseinzadeh and Y. Wang, “Image Change Captioning by Learning from an Auxiliary Task,” in CVPR , 2021
2021
Closest in time.
F. Liu, X. Wu, S. Ge, W. Fan, and Y. Zou, “Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation,” in CVPR , 2021
2021
Closest in time.
X. Yang, M. Ye, Q. You, and F. Ma, “Writing by Memorizing: Hierarchical Retrieval-based Medical Report Generation,” in ACL-IJCNLP , 2021
2021
Closest in time.
Z. Bai, Y. Nakashima, and N. Garcia, “Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation,” in ICCV , 2021
2021
Closest in time.
F. Liu, Y. Wang, T. Wang, and V. Ordonez, “Visual News: Benchmark and Challenges in News Image Captioning,” in EMNLP , 2021
2021
Closest in time.
2021
Closest in time.
Z. Meng, L. Yu, N. Zhang, T. Berg, B. Damavandi, V. Singh, and A. Bearman, “Connecting What to Say With Where to Look by Modeling Human Attention Traces,” in CVPR , 2021
2021
Closest in time.
L. Chen, Z. Jiang, J. Xiao, and W. Liu, “Human-like Controllable Image Captioning with Verb-specific Semantic Roles,” in CVPR , 2021
2021
Closest in time.