Fetching the paper…
Reading the bibliography…
Vision--Language Models (VLMs) have demonstrated success across diverse applications, yet their potential to assist in relevance judgments remains uncertain.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
The ASLIB cranfield research project on the comparative efficiency of indexing systems
Cyril W Cleverdon. 1960 · 1960
Earlier work this paper cites.
Variations in relevance judgments and the measurement of retrieval effectiveness
Ellen M Voorhees. 1998 · 1998
Earlier work this paper cites.
Humans optional? Automatic large-scale test collections for entity, passage, and entity-passage retrieval
Laura Dietz and Jeff Dalton. 2020 · 2020
Earlier work this paper cites.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Benchmark for compositional text-to-image synthesis
Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Earlier work this paper cites.
Towards multi-modal text-image retrieval to improve human reading
Florian Schneider, Özge Alaçam, Xintong Wang, and Chris Biemann. 2021 · 2021
Earlier work this paper cites.
Mutual information divergence: A unified metric for multimodal generative models
Jin-Hwa Kim, Yunji Kim, Jiyoung Lee, Kang Min Yoo, and Sang-Woo Lee. 2022 · 2022
Earlier work this paper cites.
Context matters for image descriptions for accessibility: Challenges for referenceless evaluation metrics
Elisa Kreiss, Cynthia Bennett, Shayan Hooshmand, Eric Zelikman, Meredith Ringel Morris, and Christopher Potts. 2022 · 2022
Earlier work this paper cites.
Let’s ViCE! Mimicking human cognitive behavior in image generation evaluation
Federico Betti, Jacopo Staiano, Lorenzo Baraldi, Rita Cucchiara, and Niculae Sebe. 2023 · 2023
Cited alongside, same era.
IC3: Image captioning by committee consensus
David Chan, Austin Myers, Sudheendra Vijayanarasimhan, David Ross, and John Canny. 2023a · 2023
Cited alongside, same era.
CLAIR: Evaluating image captions with large language models
David M Chan, Suzanne Petryk, Joseph E Gonzalez, Trevor Darrell, and John Canny. 2023b · 2023
Cited alongside, same era.
LLMs may dominate information access: Neural retrievers are biased towards LLM-generated texts
Sunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu, Xiaolin Hu, Yong Liu, Xiao Zhang, and Jun Xu. 2023 · 2023
Cited alongside, same era.
Perspectives on large language models for relevance judgment
Guglielmo Faggioli, Laura Dietz, Charles LA Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, et al. 2023 · 2023
Cited alongside, same era.
GPT-4V(ision) system card
OpenAI. 2023 · 2023
Later among the works it cites.
Automated annotation with generative ai requires validation
Nicholas Pangakis, Samuel Wolken, and Neil Fasching. 2023 · 2023
Later among the works it cites.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, et al. 2023 · 2023
Later among the works it cites.
DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023 · 2023
Later among the works it cites.
Enhancing textbooks with visuals from the web for improved learning
Janvijay Singh, Vilém Zouhar, and Mrinmaya Sachan. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TIFA: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A. Smith. 2023 · 2023
Cited alongside, same era.
ContextRef: Evaluating referenceless metrics for image description generation
Elisa Kreiss*, Eric Zelikman*, Christopher Potts, and Nick Haber. 2023 · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a · 2023
Cited alongside, same era.
LLMScore: Unveiling the power of large language models in text-to-image synthesis evaluation
Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. 2023 · 2023
Cited alongside, same era.
One-shot labeling for automatic relevance estimation
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b
Cited in the paper.
GPTEval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023c
Cited in the paper.
Is ChatGPT good at search? Investigating large language models as re-ranking agents
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023 · 2023
Later among the works it cites.
AToMiC: An image/text retrieval test collection to support multimedia content creation
Jheng-Hong Yang, Carlos Lassance, Rafael Sampaio De Rezende, Krishna Srinivasan, Miriam Redi, Stéphane Clinchant, and Jimmy Lin. 2023 · 2023
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. 2023 · 2023
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence. 2023 · 2023
Later among the works it cites.
Beyond yes and no: Improving zero-shot LLM rankers via scoring fine-grained relevance labels
Honglei Zhuang, Zhen Qin, Kai Hui, Junru Wu, Le Yan, Xuanhui Wang, and Michael Berdersky. 2023 · 2023
Later among the works it cites.