Fetching the paper…
Reading the bibliography…
Posters play a crucial role in marketing and advertising by enhancing visual communication and brand visibility, making significant contributions to industrial design.
ICDAR 2013 Robust Reading Competition
Karatzas, D.; Shafait, F.; Uchida, S.; Iwamura, M.; Bigorda, L. G. i.; Mestre, S. R.; Mas, J.; Mota, D. F.; Almazàn, J. A.; and de las Heras, L. P. 2013 · 2013
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Earlier work this paper cites.
ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
Sun, Z.; Li, X.; Sun, X.; Meng, Y.; Ao, X.; He, Q.; Wu, F.; and Li, J. 2021 · 2021
Earlier work this paper cites.
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
Gu, J.; Meng, X.; Lu, G.; Hou, L.; Minzhe, N.; Liang, X.; Yao, L.; Huang, R.; Zhang, W.; Jiang, X.; et al. 2022 · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Earlier work this paper cites.
Resolution-robust large mask inpainting with fourier convolutions
Suvorov, R.; Logacheva, E.; Mashikhin, A.; Remizova, A.; Ashukha, A.; Silvestrov, A.; Kong, N.; Goka, H.; Park, K.; and Lempitsky, V. 2022 · 2022
Earlier work this paper cites.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L.; Barua, A.; Constant, N.; Al-Rfou, R.; Narang, S.; Kale, M.; Roberts, A.; and Raffel, C. 2022 · 2022
Earlier work this paper cites.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; Hui, B.; Ji, L.; Li, M.; Lin, J.; Lin, R.; Liu, D.; Liu, G.; Lu, C.; Lu, K.; Ma, J.; Men, R.; Ren, X.; Ren, X.; Tan, C.; Tan, S.; Tu, J.; Wang, P.; Wang, S.; Wang, W.; Wu, S.; Xu, B.; Xu, J.; Yang, A.; Yang, H.; Yang, J.; Yang, S.; Yao, Y.; Yu, B.; Yuan, H.; Yuan, Z.; Zhang, J.; Zhang, X.; Zhang, Y.; Zhang, Z.; Zhou, C.; Zhou, J.; Zhou, X.; and Zhu, T. 2023 · 2023
Earlier work this paper cites.
Llm blueprint: Enabling text-to-image generation with complex and detailed prompts
Gani, H.; Bhat, S. F.; Naseer, M.; Khan, S.; and Wonka, P. 2023 · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Unicontrol: A unified diffusion model for controllable visual generation in the wild
Qin, C.; Zhang, S.; Yu, N.; Feng, Y.; Yang, X.; Zhou, Y.; Wang, H.; Niebles, J. C.; Xiong, C.; Savarese, S.; et al. 2023 · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Layoutgpt: Compositional visual planning and generation with large language models
Feng, W.; Zhu, W.; Fu, T.-j.; Jampani, V.; Akula, A.; He, X.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2024 · 2024
Closest in time.
SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-form Layout-to-Image Generation
Jia, C.; Luo, M.; Dang, Z.; Dai, G.; Chang, X.; Wang, M.; and Wang, J. 2024 · 2024
Closest in time.
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
Lakhanpal, S.; Chopra, S.; Jain, V.; Chadha, A.; and Luo, M. 2024 · 2024
Closest in time.
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing
Li, D.; Li, J.; and Hoi, S. 2024 · 2024
Closest in time.
LayoutPrompter: Awaken the Design Ability of Large Language Models
Lin, J.; Guo, J.; Sun, S.; Yang, Z.; Lou, J.-G.; and Zhang, D. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
AnyText: Multilingual Visual Text Generation And Editing
Tuo, Y.; Xiang, W.; He, J.-Y.; Geng, Y.; and Xie, X. 2023 · 2023
Cited alongside, same era.
Wu, X.; Hao, Y.; Sun, K.; Chen, Y.; Zhu, F.; Zhao, R.; and Li, H. 2023 · 2023
Cited alongside, same era.
Baichuan 2: Open Large-scale Language Models
Yang, A.; Xiao, B.; Wang, B.; Zhang, B.; Bian, C.; Yin, C.; Lv, C.; Pan, D.; Wang, D.; Yan, D.; Yang, F.; Deng, F.; Wang, F.; Liu, F.; Ai, G.; Dong, G.; Zhao, H.; Xu, H.; Sun, H.; Zhang, H.; Liu, H.; Ji, J.; Xie, J.; Dai, J.; Fang, K.; Su, L.; Song, L.; Liu, L.; Ru, L.; Ma, L.; Wang, M.; Liu, M.; Lin, M.; Nie, N.; Guo, P.; Sun, R.; Zhang, T.; Li, T.; Li, T.; Cheng, W.; Chen, W.; Zeng, X.; Wang, X.; Chen, X.; Men, X.; Yu, X.; Pan, X.; Shen, Y.; Wang, Y.; Li, Y.; Jiang, Y.; Gao, Y.; Zhang, Y.; Zhou, Z.; and Wu, Z. 2023 · 2023
Cited alongside, same era.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H.; Zhang, J.; Liu, S.; Han, X.; and Yang, W. 2023 · 2023
Cited alongside, same era.
Points-to-3d: Bridging the gap between sparse points and shape-controllable text-to-3d generation
Yu, C.; Zhou, Q.; Li, J.; Zhang, Z.; Wang, Z.; and Wang, F. 2023 · 2023
Cited alongside, same era.
Zavadski, D.; Feiden, J.-F.; and Rother, C. 2023 · 2023
Cited alongside, same era.
Zhao, Y.; and Lian, Z. 2023 · 2023
Cited alongside, same era.
Textdiffuser: Diffusion models as text painters
Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; and Wei, F. 2024 · 2024
Cited alongside, same era.
Closest in time.
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
Liu, Z.; Liang, W.; Liang, Z.; Luo, C.; Li, J.; Huang, G.; and Yuan, Y. 2024 · 2024
Closest in time.
Compositional Text-to-Image Generation with Dense Blob Representations
Nie, W.; Liu, S.; Mardani, M.; Liu, C.; Eckart, B.; and Vahdat, A. 2024 · 2024
Closest in time.
CustomText: Customized Textual Image Generation using Diffusion Models
Paliwal, S.; Jain, A.; Sharma, M.; Jamwal, V.; and Vig, L. 2024 · 2024
Closest in time.
Kolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis
Team, K. 2024 · 2024
Closest in time.
GlyphControl: Glyph Conditional Control for Visual Text Generation
Yang, Y.; Gui, D.; Yuan, Y.; Liang, W.; Ding, H.; Hu, H.; and Chen, K. 2024 · 2024
Closest in time.
EAFormer: Scene Text Segmentation with Edge-Aware Transformers
Yu, H.; Fu, T.; Li, B.; and Xue, X. 2024 · 2024
Closest in time.
ARTIST: Improving the Generation of Text-rich Images by Disentanglement
Zhang, J.; Zhou, Y.; Gu, J.; Wigington, C.; Yu, T.; Chen, Y.; Sun, T.; and Zhang, R. 2024 · 2024
Closest in time.
Layout-Agnostic Scene Text Image Synthesis with Diffusion Models
Zhangli, Q.; Jiang, J.; Liu, D.; Yu, L.; Dai, X.; Ramchandani, A.; Pang, G.; Metaxas, D. N.; and Krishnan, P. 2024 · 2024
Closest in time.