Fetching the paper…
Reading the bibliography…
We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model.
Making the v in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
WikiHow: A large scale text summarization dataset
M. Koupaee and W. Y. Wang · 2018
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
KVQA: Knowledge-aware visual question answering
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, et al · 2020
Earlier work this paper cites.
Diffusion models beat GANs on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
IconQA: A new benchmark for abstract diagram understanding and visual language reasoning
P. Lu, L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Earlier work this paper cites.
WIT: Wikipedia-based image text dataset for multimodal multilingual machine learning
K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
LAION-Aesthetics-UMAP, 2022
dclure · 2022
Earlier work this paper cites.
ScreenQA: Large-scale question-answer pairs over mobile app screenshots
Y.-C. Hsiao, F. Zubach, G. Baechler, V. Carbune, J. Lin, M. Wang, S. Sunkara, Y. Zhu, and J. Chen · 2022
Earlier work this paper cites.
Rectified flow: A marginal preserving approach to optimal transport
Q. Liu · 2022
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with CLIP latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation
L. Ruan, Y. Ma, H. Yang, H. He, B. Liu, J. Fu, N. J. Yuan, Q. Jin, and B. Guo · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Building normalizing flows with stochastic interpolants
M. Albergo and E. Vanden-Eijnden · 2023
Earlier work this paper cites.
Qwen-VL: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
All are worth words: A ViT backbone for diffusion models
F. Bao, S. Nie, K. Xue, Y. Cao, C. Li, H. Su, and J. Zhu · 2023
Cited alongside, same era.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Cited alongside, same era.
MobileVLM: A fast, reproducible and strong vision language assistant for mobile devices
X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei, et al · 2023
MME: A comprehensive evaluation benchmark for multimodal large language models
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji · 2024
Closest in time.
SEED-X: Multimodal models with unified multi-granularity comprehension and generation
Y. Ge, S. Zhao, J. Zhu, Y. Ge, K. Yi, L. Song, C. Li, X. Ding, and Y. Shan · 2024
Closest in time.
GenEval: An object-focused framework for evaluating text-to-image alignment
D. Ghosh, H. Hajishirzi, and L. Schmidt · 2024
Closest in time.
ELLA: Equip diffusion models with llm for enhanced semantic alignment
X. Hu, R. Wang, Y. Fang, B. Fu, P. Cheng, and G. Yu · 2024
Closest in time.
AlphaFold meets flow matching for generating protein ensembles
B. Jing, B. Berger, and T. Jaakkola · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
DeepFloyd IF, 2023
DeepFloyd · 2023
Cited alongside, same era.
Detailed caption, 2023
echo840 · 2023
Cited alongside, same era.
HAI-LLM: Efficient and lightweight training tool for large models, 2023
High-flyer · 2023
Cited alongside, same era.
Scaling up GANs for text-to-image synthesis
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park · 2023
Cited alongside, same era.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Cited alongside, same era.
Introducing IDEFICS: An open reproduction of state-of-the-art visual language model, 2023, 2023
H. Laurençon, D. van Strien, S. Bekman, L. Tronchon, L. Saulnier, T. Wang, S. Karamcheti, A. Singh, G. Pistilli, Y. Jernite, et al · 2023
Cited alongside, same era.
Closest in time.
P-Flow: a fast and data-efficient zero-shot tts through speech prompting
S. Kim, K. Shih, J. F. Santos, E. Bakhturina, M. Desta, R. Valle, S. Yoon, B. Catanzaro, et al · 2024
Closest in time.
VoiceBox: Text-guided multilingual universal speech generation at scale
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, et al · 2024
Closest in time.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge, 2024b
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
DeepSeek-VL: towards real-world vision-language understanding
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al · 2024
Closest in time.
SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie · 2024
Closest in time.
Megalith-10M, 2024
madebyollin · 2024
Closest in time.
YFCC-15M, 2024
mehdidc · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2024
Closest in time.
PyTorch, 2024
PyTorch-Contributors · 2024
Closest in time.
From pixels to prose: A large dataset of dense image captions
V. Singla, K. Yue, S. Paul, R. Shirkavand, M. Jayawardhana, A. Ganjdanesh, H. Huang, A. Bhatele, G. Somepalli, and T. Goldstein · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
C. Team · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, et al · 2024
Closest in time.
Greedy growing enables high-resolution pixel-based diffusion models
C. N. Vasconcelos, A. Rashwan, A. Waters, T. Walker, K. Xu, J. Yan, R. Qian, Y. Li, S. LUO, Y. Onoe, et al · 2024
Closest in time.
Show-o: One single transformer to unify multimodal understanding and generation
J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, and M. Z. Shou · 2024
Closest in time.
X-VILA: Cross-modality alignment for large language model
H. Ye, D.-A. Huang, Y. Lu, Z. Yu, W. Ping, A. Tao, J. Kautz, S. Han, D. Xu, P. Molchanov, et al · 2024
Closest in time.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Closest in time.
MonoFormer: One transformer for both diffusion and autoregression
C. Zhao, Y. Song, W. Wang, H. Feng, E. Ding, Y. Sun, X. Xiao, and J. Wang · 2024
Closest in time.
Transfusion: Predict the next token and diffuse images with one multi-modal model
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy · 2024
Closest in time.
LLaVA-Phi: Efficient multi-modal assistant with small language model
Y. Zhu, M. Zhu, N. Liu, Z. Ou, X. Mou, and J. Tang · 2024
Closest in time.
Lumina-Next: Making Lumina-T2X stronger and faster with Next-DiT
L. Zhuo, R. Du, H. Xiao, Y. Li, D. Liu, R. Huang, W. Liu, L. Zhao, F.-Y. Wang, Z. Ma, et al · 2024
Closest in time.