Fetching the paper…
Reading the bibliography…
Recent image generation models excel at creating high-quality images from brief captions.
“Bleu: a method for automatic evaluation of machine translation”
Kishore Papineni, Salim Roukos, Todd Ward and Wei-Jing Zhu · 2002
Earlier work this paper cites.
“NLTK: the natural language toolkit”
Steven Bird · 2006
Earlier work this paper cites.
“Visual Storytelling” arXiv:1604.03968 [cs]
Ting-Hao et al · 2016
Earlier work this paper cites.
“StoryGAN: A Sequential Conditional GAN for Story Visualization”, 2019
Yitong Li et al · 2019
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“Integrating Visuospatial, Linguistic and Commonsense Structure into Story Visualization”, 2021
Adyasha Maharana and Mohit Bansal · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“Laion-400m: Open dataset of clip-filtered 400 million image-text pairs”
Christoph Schuhmann et al · 2021
Earlier work this paper cites.
“GLM: General Language Model Pretraining with Autoregressive Blank Infilling”
Zhengxiao Du et al · 2022
Earlier work this paper cites.
“StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation”, 2022
Adyasha Maharana, Darryl Hannan and Mohit Bansal · 2022
Earlier work this paper cites.
“High-Resolution Image Synthesis with Latent Diffusion Models”, 2022
Robin Rombach et al · 2022
Earlier work this paper cites.
“High-resolution image synthesis with latent diffusion models”
Robin Rombach et al · 2022
Earlier work this paper cites.
“Glm-130b: An open bilingual pre-trained model”
Aohan Zeng et al · 2022
Earlier work this paper cites.
Josh Achiam et al · 2023
Earlier work this paper cites.
“Touchstone: Evaluating vision-language models by language models”
Shuai Bai et al · 2023
Earlier work this paper cites.
“Improving Image Generation with Better Captions”, 2023
James Betker et al · 2023
Earlier work this paper cites.
“Dreamllm: Synergistic multimodal comprehension and creation”
Runpei Dong et al · 2023
Earlier work this paper cites.
“MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models”
Chaoyou Fu et al · 2023
Cited alongside, same era.
“Making llama see and draw with seed tokenizer”
Yuying Ge et al · 2023
Cited alongside, same era.
Alexander Kirillov et al · 2023
Cited alongside, same era.
“Seed-bench-2: Benchmarking multimodal large language models”
Bohao Li et al · 2023
Cited alongside, same era.
“mplug-owl: Modularization empowers large language models with multimodality”
Qinghao Ye et al · 2023
Later among the works it cites.
“Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi”
Xiang Yue et al · 2023
Later among the works it cites.
“Minigpt-5: Interleaved vision-and-language generation via generative vokens”
Kaizhi Zheng, Xuehai He and Xin Wang · 2023
Later among the works it cites.
“Minigpt-4: Enhancing vision-language understanding with advanced large language models”
Deyao Zhu et al · 2023
Later among the works it cites.
“Phi-3 technical report: A highly capable language model locally on your phone”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Cited alongside, same era.
“Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models”
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Cited alongside, same era.
“Video-LLaVA: Learning United Visual Representation by Alignment Before Projection”
Bin Lin et al · 2023
Cited alongside, same era.
“VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning”, 2023
Han Lin, Abhay Zala, Jaemin Cho and Mohit Bansal · 2023
Cited alongside, same era.
“Mmbench: Is your multi-modal model an all-around player?”
Yuan Liu et al · 2023
Cited alongside, same era.
Jian Ma, Junhao Liang, Chen Chen and Haonan Lu · 2023
Cited alongside, same era.
“DINOv2: Learning Robust Visual Features without Supervision”
Maxime Oquab et al · 2023
Cited alongside, same era.
“Kosmos-g: Generating images in context with multimodal large language models”
Xichen Pan et al · 2023
Cited alongside, same era.
Marah Abdin et al · 2024
Closest in time.
“Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers”
Tsai-Shien Chen et al · 2024
Closest in time.
“YOLO-World: Real-Time Open-Vocabulary Object Detection”
Tianheng Cheng et al · 2024
Closest in time.
“LAMM: Label Alignment for Multi-Modal Prompt Learning”
Jingsheng Gao et al · 2024
Closest in time.
“Seed-x: Multimodal models with unified multi-granularity comprehension and generation”
Yuying Ge et al · 2024
Closest in time.
“Seed-x: Multimodal models with unified multi-granularity comprehension and generation”
Yuying Ge et al · 2024
Closest in time.
“Instruct-Imagen: Image generation with multi-modal instruction”
Hexiang Hu et al · 2024
Closest in time.
“Mini-gemini: Mining the potential of multi-modality vision language models”
Yanwei Li et al · 2024
Closest in time.
Chang Liu et al · 2024
Closest in time.
“Visual instruction tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2024
Closest in time.
“Journeydb: A benchmark for generative image understanding”
Keqiang Sun et al · 2024
Closest in time.
“EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss”
Zhuoyang Zhang, Han Cai and Song Han · 2024
Closest in time.
“Judging llm-as-a-judge with mt-bench and chatbot arena”
Lianmin Zheng et al · 2024
Closest in time.