Fetching the paper…
Reading the bibliography…
Referenceless metrics (e.g., CLIPScore) use pretrained vision--language models to assess image descriptions directly without costly ground-truth reference texts.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Catching the drift: Probabilistic content models, with applications to generation and summarization
Regina Barzilay and Lillian Lee · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Jesse Dodge, Amit Goyal, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa Mensch, Alexander Berg, Tamara Berg, and Hal Daumé III · 2012
Earlier work this paper cites.
Image description using visual dependency representations
Desmond Elliott and Frank Keller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Zero-shot learning by convex combination of semantic embeddings
Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S Corrado, and Jeffrey Dean · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2017
Earlier work this paper cites.
Room for improvement in automatic image description: an error analysis
Emiel van Miltenburg and Desmond Elliott · 2017
Earlier work this paper cites.
Learning to evaluate image captioning
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge Belongie · 2018
Earlier work this paper cites.
Prolific. ac—a subject pool for online experiments
Stefan Palan and Christian Schitter · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko · 2018
Earlier work this paper cites.
“it’s almost like they’re trying to hide it”: How user-provided image descriptions have failed to make twitter accessible
Cole Gleason, Patrick Carrington, Cameron Cassidy, Meredith Ringel Morris, Kris M Kitani, and Jeffrey P Bigham · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Unsupervised parsing via constituency tests
Steven Cao, Nikita Kitaev, and Dan Klein · 2020
Cited alongside, same era.
Pragmatic issue-sensitive image captioning
Allen Nie, Reuben Cohn-Gordon, and Christopher Potts · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Cited alongside, same era.
" person, shoes, tree. is the person naked?" what people with vision impairments want in image descriptions
Abigale Stangl, Meredith Ringel Morris, and Danna Gurari · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Concadia: Towards image-based text generation with a purpose
Elisa Kreiss, Fei Fang, Noah Goodman, and Christopher Potts · 2022
Later among the works it cites.
What’s in an alt tag? exploring caption content priorities through collaborative captioning
Annika Muehlbradt and Shaun K Kane · 2022
Later among the works it cites.
Are multimodal models robust to image and text perturbations?
Jielin Qiu, Yi Zhu, Xingjian Shi, Florian Wenzel, Zhiqiang Tang, Ding Zhao, Bo Li, and Mu Li · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Toward supporting quality alt text in computing publications
Candace Williams, Lilian de Greef, Ed Harris III, Leah Findlater, Amy Pavel, and Cynthia Bennett · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
OpenCLIP
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Cited alongside, same era.
Qace: Asking questions to evaluate an image caption
Hwanhee Lee, Thomas Scialom, Seunghyun Yoon, Franck Dernoncourt, and Kyomin Jung · 2021
Cited alongside, same era.
Sometimes we want ungrammatical translations
Prasanna Parthasarathi, Koustuv Sinha, Joelle Pineau, and Adina Williams · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, A. Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela · 2021
Cited alongside, same era.
Learning to break the loop: Analyzing and mitigating repetitions for neural text generation
Jin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai, Huayang Li, and Jian Li · 2022
Later among the works it cites.
Openflamingo
Anas Awadalla, Irena Gao, Joshua Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt · 2023
Closest in time.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Closest in time.
Data quality in online human-subjects research: Comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA
Benjamin D Douglas, Patrick J Ewell, and Markus Brauer · 2023
Closest in time.
Multimodal few-shot learning by convex combination of token embeddings
dzryk · 2023
Closest in time.
Eva-02: A visual representation for neon genesis
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2023
Closest in time.
PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning
Yongil Kim, Yerin Hwang, Hyeongu Yun, Seunghyun Yoon, Trung Bui, and Kyomin Jung · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Visualgptscore: Visio-linguistic reasoning with multimodal generative pre-training scores
Zhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang, and Deva Ramanan · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Dealing with semantic underspecification in multimodal nlp
Sandro Pezzelle · 2023
Closest in time.
Improved image caption rating–datasets, game, and model
Andrew Taylor Scott, Lothar D Narins, Anagha Kulkarni, Mar Castanon, Benjamin Kao, Shasta Ihorn, Yue-Ting Siu, and Ilmi Yoon · 2023
Closest in time.
Vistext: A benchmark for semantically rich chart captioning
Benny J Tang, Angie Boggust, and Arvind Satyanarayan · 2023
Closest in time.