Fetching the paper…
Reading the bibliography…
Generating images with accurately represented text, especially in non-Latin languages, poses a significant challenge for diffusion models.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.-Y.; and Ermon, S. 2021 · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
eDiff-I: Text-to-Image Diffusion Models with Ensemble of Expert Denoisers
Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Zhang, Q.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; Catanzaro, B.; Karras, T.; and Liu, M.-Y. 2022 · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Earlier work this paper cites.
Caponimage: Context-driven dense-captioning on image
Gao, Y.; Hou, X.; Zhang, Y.; Ge, T.; Jiang, Y.; and Wang, P. 2022 · 2022
Earlier work this paper cites.
Scalable Diffusion Models with Transformers
Peebles, W.; and Xie, S. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Cited alongside, same era.
Deepfloyd if
DeepFloyd-Lab. 2023 · 2023
Cited alongside, same era.
Composer: Creative and controllable image synthesis with composable conditions
Cogvlm: Visual expert for pretrained language models
Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. 2023 · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H.; Zhang, J.; Liu, S.; Han, X.; and Yang, W. 2023 · 2023
Later among the works it cites.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
Zhao, Y.; and Lian, Z. 2023 · 2023
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, L.; Chen, D.; Liu, Y.; Shen, Y.; Zhao, D.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Imagic: Text-based real image editing with diffusion models
Kawar, B.; Zada, S.; Lang, O.; Tov, O.; Chang, H.; Dekel, T.; Mosseri, I.; and Irani, M. 2023 · 2023
Cited alongside, same era.
Glyphdraw: Learning to draw chinese characters in image synthesis models coherently
Ma, J.; Zhao, M.; Chen, C.; Wang, R.; Niu, D.; Lu, H.; and Lin, X. 2023 · 2023
Cited alongside, same era.
Duguangocr
ModelScope. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
AnyText: Multilingual Visual Text Generation and Editing
Tuo, Y.; Xiang, W.; He, J.-Y.; Geng, Y.; and Xie, X. 2023 · 2023
Cited alongside, same era.
TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering
Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; and Wei, F. 2023a
Cited in the paper.
Closest in time.
Glyph-byt5: A customized text encoder for accurate visual text rendering
Liu, Z.; Liang, W.; Liang, Z.; Luo, C.; Li, J.; Huang, G.; and Yuan, Y. 2024 · 2024
Closest in time.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Mou, C.; Wang, X.; Xie, L.; Wu, Y.; Zhang, J.; Qi, Z.; and Shan, Y. 2024 · 2024
Closest in time.
GlyphControl: Glyph Conditional Control for Visual Text Generation
Yang, Y.; Gui, D.; Yuan, Y.; Liang, W.; Ding, H.; Hu, H.; and Chen, K. 2024 · 2024
Closest in time.
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
Ye, M.; Zhang, J.; Liu, J.; Liu, C.; Yin, B.; Liu, C.; Du, B.; and Tao, D. 2024 · 2024
Closest in time.
Brush your text: Synthesize any scene text on images via diffusion model
Zhang, L.; Chen, X.; Wang, Y.; Lu, Y.; and Qiao, Y. 2024 · 2024
Closest in time.