Fetching the paper…
Reading the bibliography…
Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Viziometrics: Analyzing visual information in the scientific literature
Po-Shen Lee, Jevin D. West, and Bill Howe · 2016
Earlier work this paper cites.
Figurefirst: A layout-first approach for scientific figures
Theodore Lindsay, Peter Weir, and Floris van Breugel · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
An automatic graph generation method for scholarly papers based on table structure analysis
Ryoya Yamada, Manabu Ohta, and Atsuhiro Takasu · 2018
Earlier work this paper cites.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Earlier work this paper cites.
seaborn: statistical data visualization
Michael Waskom · 2021
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Earlier work this paper cites.
Elicit: Language models as research tools
Jungwon Byun and Andreas Stuhlmüller · 2023
Cited alongside, same era.
Transformers go for the LOLs: Generating (humourous) titles from scientific abstracts end-to-end
Yanran Chen and Steffen Eger · 2023
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
Jaemin Cho, Abhay Zala, and Mohit Bansal · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu · 2023
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation, 2023
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy · 2023
Cited alongside, same era.
State of What Art? A Call for Multi-Prompt LLM Evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky · 2024
Closest in time.
Introducing openai o1 (preview)
OpenAI · 2024
Closest in time.
Scifibench: Benchmarking large multimodal models for scientific figure interpretation
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie · 2024
Closest in time.
Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu, Aditya Sharma, and William Yang Wang · 2024
Closest in time.
Assisting in writing wikipedia-like articles from scratch with large language models, 2024
Yijia Shao, Yucheng Jiang, Theodore A. Kanell, Peter Xu, Omar Khattab, and Monica S. Lam · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Eval4NLP 2023 shared task on prompting large language models as explainable metrics
Christoph Leiter, Juri Opitz, Daniel Deutsch, Yang Gao, Rotem Dror, and Steffen Eger · 2023
Cited alongside, same era.
Ocr-vqgan: Taming text-within-image generation
Juan A Rodriguez, David Vazquez, Issam Laradji, Marco Pedersoli, and Pau Rodriguez · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2023
Cited alongside, same era.
Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation
Jaemin Cho, Yushi Hu, Jason Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang · 2024
Cited alongside, same era.
Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models
Rocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Iliyanov Dimitrov, Ivan Koychev, and Preslav Nakov · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Cited alongside, same era.
Closest in time.
Chartmimic: Evaluating lmm’s cross-modal reasoning capability via chart-to-code generation
Chufan Shi, Cheng Yang, Yaxin Liu, Bo Shui, Junjie Wang, Mohan Jing, Linran Xu, Xinyu Zhu, Siheng Li, Yuxiang Zhang, et al · 2024
Closest in time.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers, 2024
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Closest in time.
Plots made quickly: An efficient approach for generating visualizations from natural language queries
Henrik Voigt, Kai Lawonn, and Sina Zarrieß · 2024
Closest in time.
Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models
Rohan Wadhawan, Hritik Bansal, Kai-Wei Chang, and Nanyun Peng · 2024
Closest in time.
Evaluating and analyzing relationship hallucinations in large vision-language models
Mingrui Wu, Jiayi Ji, Oucheng Huang, Jiale Li, Yuhang Wu, Xiaoshuai Sun, and Rongrong Ji · 2024
Closest in time.
Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi
Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, et al · 2024
Closest in time.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang · 2024
Closest in time.
Vgbench: Evaluating large language models on vector graphics understanding and generation
Bocheng Zou, Mu Cai, Jianrui Zhang, and Yong Jae Lee · 2024
Closest in time.
Evaluating large language models for structured science summarization in the open research knowledge graph
Vladyslav Nechakhin, Jennifer D’Souza, and Steffen Eger · 2078
Closest in time.