Fetching the paper…
Reading the bibliography…
Image captioning systems are unable to generate fine-grained captions as they are trained on data that is either noisy (alt-text) or generic (human annotations).
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
Ronald J Williams · 1992
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
CIDEr: Consensus-Based Image Description Evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
SPICE: Semantic Propositional Image Caption Evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
YFCC100M: The New Data in Multimedia Research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
Abhishek Das, Satwik Kottur, José MF Moura, Stefan Lee, and Dhruv Batra · 2017
Earlier work this paper cites.
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Earlier work this paper cites.
Self-Critical Sequence Training for Image Captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel · 2017
Earlier work this paper cites.
Learning to Describe Differences Between Pairs of Similar Images
Harsh Jhamtani and Taylor Berg-Kirkpatrick · 2018
Earlier work this paper cites.
Show, Tell and Discriminate: Image Captioning by Self-retrieval with Partially Labeled Data
Xihui Liu, Hongsheng Li, Jing Shao, Dapeng Chen, and Xiaogang Wang · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Discriminability Objective for Training Descriptive Captions
Ruotian Luo, Brian Price, Scott Cohen, and Gregory Shakhnarovich · 2018
Earlier work this paper cites.
Object Hallucination in Image Captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko · 2018
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Robust Change Captioning
Dong Huk Park, Trevor Darrell, and Anna Rohrbach · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
CapWAP: Image Captioning with a Purpose
Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H Clark, and Regina Barzilay · 2020
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Towards Unique and Informative Captioning of Images
Zeyu Wang, Berthy Feng, Karthik Narasimhan, and Olga Russakovsky · 2020
Cited alongside, same era.
RedCaps: Web-curated image-text data created by the people, for the people
Karan Desai, Gaurav Kaul, Zubin Trivadi Aysola, and Justin Johnson · 2021
Cited alongside, same era.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Cited alongside, same era.
Introducing our multimodal models, 2023
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sağnak Taşırlar · 2023
Later among the works it cites.
Cross-Domain Image Captioning With Discriminative Finetuning
Roberto Dessì, Michele Bevilacqua, Eleonora Gualdoni, Nathanaël Carraz Rakotonirina, Francesca Franzon, and Marco Baroni · 2023
Later among the works it cites.
Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities
Melissa Hall, Laura Gustafson, Aaron Adcock, Ishan Misra, and Candace Ross · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
Guiding Image Captioning Models Toward More Specific Captions
Simon Kornblith, Lala Li, Zirui Wang, and Thao Nguyen · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ron Mokady, Amir Hertz, and Amit H Bermano · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
CIDEr-R: Robust Consensus-based Image Description Evaluation
Gabriel Oliveira dos Santos, Esther Luna Colombini, and Sandra Avila · 2021
Cited alongside, same era.
Big Vision
Lucas Beyer, Xiaohua Zhai, and Alexander Kolesnikov · 2022
Cited alongside, same era.
Fine-grained Image Captioning with CLIP Reward
Jaemin Cho, Seunghyun Yoon, Ajinkya Kale, Franck Dernoncourt, Trung Bui, and Mohit Bansal · 2022
Cited alongside, same era.
Concadia: Towards Image-Based Text Generation with a Purpose
Elisa Kreiss, Fei Fang, Noah Goodman, and Christopher Potts · 2022
Cited alongside, same era.
Image Retrieval from Contextual Descriptions
Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Ponti, and Siva Reddy · 2022
Cited alongside, same era.
VeCLIP: Improving CLIP Training via Visual-enriched Captions
Zhengfeng Lai, Haotian Zhang, Wentao Wu, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang, et al · 2023
Later among the works it cites.
Improved Baselines with Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2023
Later among the works it cites.
Tuning Computer Vision Models with Task Rewards
André Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer, and Xiaohua Zhai · 2023
Later among the works it cites.
Computer Vision Datasets and Models Exhibit Cultural and Linguistic Diversity in Perception
Andre Ye, Sebastin Santy, Jena D Hwang, Amy X Zhang, and Ranjay Krishna · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models
Sivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig, Donghyun Kim, Paola Cascante-Bonilla, Amit Alfassy, Rameswar Panda, Raja Giryes, Rogerio Feris, et al · 2024
Closest in time.
Describing Differences in Image Sets with Natural Language
Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E Gonzalez, and Serena Yeung-Levy · 2024
Closest in time.
Multi-modal Hallucination Control by Visual Information Grounding
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto · 2024
Closest in time.
SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna · 2024
Closest in time.
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
Qirui Jiao, Daoyuan Chen, Yilun Huang, Yaliang Li, and Ying Shen · 2024
Closest in time.
What If We Recaption Billions of Web Images with LLaMA-3?
Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, et al · 2024
Closest in time.
A Survey on Hallucination in Large Vision-Language Models
Hanchao Liu, Wenyuan Xue, Yifei Chen, Dapeng Chen, Xiutian Zhao, Ke Wang, Liping Hou, Rongjun Li, and Wei Peng · 2024
Closest in time.
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht, Bernt Schiele, and Hilde Kuehne · 2024
Closest in time.
From Pixels to Prose: A Large Dataset of Dense Image Captions
Vasu Singla, Kaiyu Yue, Sukriti Paul, Reza Shirkavand, Mayuka Jayawardhana, Alireza Ganjdanesh, Heng Huang, Abhinav Bhatele, Gowthami Somepalli, and Tom Goldstein · 2024
Closest in time.
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, and Adriana Romero-Soriano · 2024
Closest in time.