Fetching the paper…
Reading the bibliography…
Controllable image captioning is an emerging multimodal topic that aims to describe the image with natural language following human purpose, $\textit{e.g.}$, looking at the specified regions or telling in a particular text style.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Earlier work this paper cites.
Senticap: Generating image descriptions with sentiments
Alexander Mathews, Lexing Xie, and Xuming He · 2016
Earlier work this paper cites.
Deep interactive object selection
Ning Xu, Brian Price, Scott Cohen, Jimei Yang, and Thomas S Huang · 2016
Earlier work this paper cites.
Stylenet: Generating attractive visual captions with styles
Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, and Li Deng · 2017
Earlier work this paper cites.
Regional interactive image segmentation networks
JunHao Liew, Yunchao Wei, Wei Xiong, Sim-Heng Ong, and Jiashi Feng · 2017
Earlier work this paper cites.
Dense captioning with joint inference and visual context
Linjie Yang, Kevin Tang, Jianchao Yang, and Li-Jia Li · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Interactive image segmentation with latent diversity
Zhuwen Li, Qifeng Chen, and Vladlen Koltun · 2018
Earlier work this paper cites.
Iteratively trained interactive segmentation
Sabarinath Mahadevan, Paul Voigtlaender, and Bastian Leibe · 2018
Earlier work this paper cites.
Deepigeos: a deep interactive geodesic framework for medical image segmentation
Guotai Wang, Maria A Zuluaga, Wenqi Li, Rosalind Pratt, Premal A Patel, Michael Aertsen, Tom Doel, Anna L David, Jan Deprest, Sébastien Ourselin, et al · 2018
Earlier work this paper cites.
Show, control and tell: A framework for generating controllable and grounded captions
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2019
Earlier work this paper cites.
Attention on attention for image captioning
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei · 2019
Earlier work this paper cites.
Learning object context for dense captioning
Xiangyang Li, Shuqiang Jiang, and Jungong Han · 2019
Earlier work this paper cites.
Context and attribute grounded dense captioning
Guojun Yin, Lu Sheng, Bin Liu, Nenghai Yu, Xiaogang Wang, and Jing Shao · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Length-controllable image captioning
Chaorui Deng, Ning Ding, Mingkui Tan, and Qi Wu · 2020
Cited alongside, same era.
Interactive image segmentation with first click attention
Zheng Lin, Zhao Zhang, Lin-Zhuo Chen, Ming-Ming Cheng, and Shao-Ping Lu · 2020
Cited alongside, same era.
Connecting vision and language with localized narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari · 2020
Cited alongside, same era.
f-brs: Rethinking backpropagating refinement for interactive segmentation
Konstantin Sofiiuk, Ilia Petrov, Olga Barinova, and Anton Konushin · 2020
Cited alongside, same era.
Memcap: Memorizing style knowledge for image captioning
Wentian Zhao, Xinxiao Wu, and Xiaoxun Zhang · 2020
Cited alongside, same era.
Human-like controllable image captioning with verb-specific semantic roles
Long Chen, Zhihong Jiang, Jun Xiao, and Wei Liu · 2021
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al · 2022
Later among the works it cites.
Region-object relation-aware dense captioning via transformer
Zhuang Shao, Jungong Han, Demetris Marnerides, and Kurt Debattista · 2022
Later among the works it cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic · 2022
Later among the works it cites.
Git: A generative image-to-text transformer for vision and language
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
https://www.jaided.ai/easyocr, 2021
EasyOCR · 2021
Cited alongside, same era.
Control image captioning spatially and temporally
Kun Yan, Lei Ji, Huaishao Luo, Ming Zhou, Nan Duan, and Shuai Ma · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
Injecting semantic concepts into end-to-end image captioning
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang, Zhe Gan, Lijuan Wang, Yezhou Yang, and Zicheng Liu · 2022
Cited alongside, same era.
Deecap: dynamic early exiting for efficient image captioning
Zhengcong Fei, Xu Yan, Shuhui Wang, and Qi Tian · 2022
Cited alongside, same era.
Scaling up vision-language pre-training for image captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Cited alongside, same era.
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Later among the works it cites.
Controllable image captioning via prompting
Ning Wang, Jiahao Xie, Jihao Wu, Mingbo Jia, and Linlin Li · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Grit: A generative region-to-text transformer for object understanding
Jialian Wu, Jianfeng Wang, Zhengyuan Yang, Zhe Gan, Zicheng Liu, Junsong Yuan, and Lijuan Wang · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Accelerating vision-language pretraining with free language modeling
Teng Wang, Yixiao Ge, Feng Zheng, Ran Cheng, Ying Shan, Xiaohu Qie, and Ping Luo · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Closest in time.
Conzic: Controllable zero-shot image captioning by sampling-based polishing
Zequn Zeng, Hao Zhang, Zhengjue Wang, Ruiying Lu, Dongsheng Wang, and Bo Chen · 2023
Closest in time.