Fetching the paper…
Reading the bibliography…
In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin · 2018
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset
M. Koupaee and W. Y. Wang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Kvqa: Knowledge-aware visual question answering
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, T. Duerig, and V. Ferrari · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
P. Lu, L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Earlier work this paper cites.
Laion-aesthetics-umap
dclure · 2022
Earlier work this paper cites.
Make-a-scene: Scene-based text-to-image generation with human priors
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Earlier work this paper cites.
Screenqa: Large-scale question-answer pairs over mobile app screenshots
Y.-C. Hsiao, F. Zubach, M. Wang, et al · 2022
Cited alongside, same era.
Beit v2: Masked image modeling with vector-quantized visual tokenizers
Z. Peng, L. Dong, H. Bao, Q. Ye, and F. Wei · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Touch and go: Learning from human-collected vision and touch
F. Yang, C. Ma, J. Zhang, J. Zhu, W. Yuan, and A. Owens · 2022
Cited alongside, same era.
Movq: Modulating quantized vectors for high-fidelity image generation
C. Zheng, T.-L. Vuong, J. Cai, and D. Phung · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Muse: Text-to-image generation via masked generative transformers
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, et al · 2023
Cited alongside, same era.
Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices
X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei, et al · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
Dreamllm: Synergistic multimodal comprehension and creation
R. Dong, C. Han, Y. Peng, Z. Qi, Z. Ge, J. Yang, L. Zhao, J. Sun, H. Zhou, H. Wei, et al · 2023
Cited alongside, same era.
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al · 2024
Closest in time.
Mobilevlm v2: Faster and stronger baseline for vision language model
X. Chu, L. Qiao, X. Zhang, S. Xu, F. Wei, Y. Yang, X. Sun, Y. Hu, X. Lin, B. Zhang, et al · 2024
Closest in time.
Seed-x: Multimodal models with unified multi-granularity comprehension and generation
Y. Ge, S. Zhao, J. Zhu, Y. Ge, K. Yi, L. Song, C. Li, X. Ding, and Y. Shan · 2024
Closest in time.
Geneval: An object-focused framework for evaluating text-to-image alignment
D. Ghosh, H. Hajishirzi, and L. Schmidt · 2024
Closest in time.
Deepseek-vl: towards real-world vision-language understanding
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, Y. Sun, et al · 2024
Closest in time.
Megalith-huggingface
madebyollin · 2024
Closest in time.
Yfcc-huggingface
mehdidc · 2024
Closest in time.
Dalle3-high-quality-captions
ProGamerGov · 2024
Closest in time.
From pixels to prose: A large dataset of dense image captions
V. Singla, K. Yue, S. Paul, R. Shirkavand, M. Jayawardhana, A. Ganjdanesh, H. Huang, A. Bhatele, G. Somepalli, and T. Goldstein · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
C. Team · 2024
Closest in time.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
K. Tian, Y. Jiang, Z. Yuan, B. Peng, and L. Wang · 2024
Closest in time.
Vila-u: a unified foundation model integrating visual understanding and generation
Y. Wu, Z. Zhang, J. Chen, H. Tang, D. Li, Y. Fang, L. Zhu, E. Xie, H. Yin, L. Yi, et al · 2024
Closest in time.
Show-o: One single transformer to unify multimodal understanding and generation
J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, and M. Z. Shou · 2024
Closest in time.
Raphael: Text-to-image generation via large mixture of diffusion paths
Z. Xue, G. Song, Q. Guo, B. Liu, Z. Zong, Y. Liu, and P. Luo · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Closest in time.
Transfusion: Predict the next token and diffuse images with one multi-modal model
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy · 2024
Closest in time.
Llava-phi: Efficient multi-modal assistant with small language model
Y. Zhu, M. Zhu, N. Liu, Z. Ou, X. Mou, and J. Tang · 2024
Closest in time.