Fetching the paper…
Reading the bibliography…
Image captioning task has been extensively researched by previous work.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Auto-encoding variational bayes, 2013
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics (extended abstract)
Micah Hodosh, Peter Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2014
Earlier work this paper cites.
Generative adversarial networks, 2014
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Knowing when to look: Adaptive attention via A visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks, 2016
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Earlier work this paper cites.
Cascade recurrent neural network for image caption generation
Jie WU and Haifeng Hu · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Image caption generation with part of speech guidance
Xinwei He, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang, and Weisheng Dong · 2019
Earlier work this paper cites.
Masked non-autoregressive image captioning
Junlong Gao, Xi Meng, Shiqi Wang, Xia Li, Shanshe Wang, Siwei Ma, and Wen Gao · 2019
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Stack-vs: Stacked visual-semantic attention for image caption generation
Wei Wei, Ling Cheng, Xianling Mao, Guangyou Zhou, and Feida Zhu · 2019
Cited alongside, same era.
Mscap: Multi-style image captioning with unpaired stylized text
Longteng Guo, Jing Liu, Peng Yao, Jiangwei Li, and Hanqing Lu · 2019
Cited alongside, same era.
Variational autoencoder-based multiple image captioning using a caption attention map
Boeun Kim, Saim Shin, and Hyedong Jung · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alex Nichol · 2021
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval, 2021
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2021
Later among the works it cites.
Actionclip: A new paradigm for video action recognition, 2021
Mengmeng Wang, Jiazheng Xing, and Yong Liu · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision, 2021
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Cited alongside, same era.
Partially non-autoregressive image captioning
Zhengcong Fei · 2021
Cited alongside, same era.
Semi-autoregressive image captioning
Xu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang, Qingming Huang, and Qi Tian · 2021
Cited alongside, same era.
GLIDE: towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2021
Cited alongside, same era.
Grad-tts: A diffusion probabilistic model for text-to-speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail A. Kudinov · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
mplug: Effective and efficient vision-language learning by cross-modal skip-connections, 2022
Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, Hehong Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, and Luo Si · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding, 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi · 2022
Closest in time.
Diffsound: Discrete diffusion model for text-to-sound generation, 2022
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu · 2022
Closest in time.
Diffusion-lm improves controllable text generation, 2022
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto · 2022
Closest in time.
Grit: Faster and better image captioning transformer using dual visual features, 2022
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani · 2022
Closest in time.
Classifier-free diffusion guidance, 2022
Jonathan Ho and Tim Salimans · 2022
Closest in time.
Language-driven semantic segmentation, 2022
Boyi Li, Kilian Q. Weinberger, Serge Belongie, Vladlen Koltun, and René Ranftl · 2022
Closest in time.
Groupvit: Semantic segmentation emerges from text supervision, 2022
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang · 2022
Closest in time.
Open-vocabulary object detection via vision and language knowledge distillation, 2022
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2022
Closest in time.
Grounded language-image pre-training, 2022
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jianfeng Gao · 2022
Closest in time.