Fetching the paper…
Reading the bibliography…
In this work, we present an unsupervised method for enhancing an image captioning model (in our case, BLIP2) using reinforcement learning and vision-language models like CLIP and BLIP2-ITM as reward models.
“Microsoft COCO: Common Objects in Context”, 2015
Tsung-Yi Lin et al · 2015
Earlier work this paper cites.
“Asynchronous Methods for Deep Reinforcement Learning”, 2016
Volodymyr Mnih et al · 2016
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”, 2017
Diederik. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
“Proximal Policy Optimization Algorithms”, 2017
John Schulman et al · 2017
Earlier work this paper cites.
Soravit Changpinyo, Piyush Sharma, Nan Ding and Radu Soricut · 2021
Earlier work this paper cites.
“OpenCLIP” If you use this software, please cite it as below
Gabriel Ilharco et al · 2021
Earlier work this paper cites.
“Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision”, 2021
Chao Jia et al · 2021
Earlier work this paper cites.
“Learning Transferable Visual Models From Natural Language Supervision”, 2021
Alec Radford et al · 2021
Earlier work this paper cites.
“LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs”, 2021
Christoph Schuhmann et al · 2021
Cited alongside, same era.
“Flamingo: a Visual Language Model for Few-Shot Learning”, 2022
Jean-Baptiste Alayrac et al · 2022
Cited alongside, same era.
Junnan Li, Dongxu Li, Caiming Xiong and Steven Hoi · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”, 2022
Long Ouyang et al · 2022
Cited alongside, same era.
“LAION-5B: An open large-scale dataset for training next generation image-text models”, 2022
“IC3: Image Captioning by Committee Consensus”, 2023
David. Chan et al · 2023
Later among the works it cites.
“Deep reinforcement learning from human preferences”, 2023
Paul Christiano et al · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Later among the works it cites.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”, 2023
Hugo Touvron et al · 2023
Later among the works it cites.
“LLaMA: Open and Efficient Foundation Language Models”, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Schuhmann et al · 2022
Cited alongside, same era.
“CoCa: Contrastive Captioners are Image-Text Foundation Models”, 2022
Jiahui Yu et al · 2022
Cited alongside, same era.
“Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language”, 2022
Andy Zeng et al · 2022
Cited alongside, same era.
“OPT: Open Pre-trained Transformer Language Models”, 2022
Susan Zhang et al · 2022
Cited alongside, same era.
Hugo Touvron et al · 2023
Later among the works it cites.
“Image Captioners Are Scalable Vision Learners Too”, 2023
Michael Tschannen et al · 2023
Later among the works it cites.
“ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions”, 2023
Deyao Zhu et al · 2023
Later among the works it cites.