Fetching the paper…
Reading the bibliography…
Achieving precise word-level typography control within generated images remains a persistent challenge.
Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1931–1941
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. 2023 · 1941
Earlier work this paper cites.
Photo aesthetics ranking network with attributes and content adaptation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14 . Springer, 662–679
Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless Fowlkes. 2016 · 2016
Earlier work this paper cites.
Awesome typography: Statistics-based text effects transfer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 7464–7473
Shuai Yang, Jiaying Liu, Zhouhui Lian, and Zongming Guo. 2017 · 2017
Earlier work this paper cites.
Multi-content gan for few-shot font style transfer. In Proceedings of the IEEE conference on computer vision and pattern recognition . 7564–7573
Samaneh Azadi, Matthew Fisher, Vladimir G Kim, Zhaowen Wang, Eli Shechtman, and Trevor Darrell. 2018 · 2018
Earlier work this paper cites.
Context-aware text-based binary image stylization and synthesis
Shuai Yang, Jiaying Liu, Wenhan Yang, and Zongming Guo. 2018a · 2018
Earlier work this paper cites.
Artistic glyph image synthesis via one-stage few-shot learning
Yue Gao, Yuan Guo, Zhouhui Lian, Yingmin Tang, and Jianguo Xiao. 2019 · 2019
Earlier work this paper cites.
Scfont: Structure-guided chinese font generation via deep stacked networks. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 4015–4022
Yue Jiang, Zhouhui Lian, Yingmin Tang, and Jianguo Xiao. 2019 · 2019
Earlier work this paper cites.
Controllable Artistic Text Style Transfer via Shape-Matching GAN. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Shuai Yang, Zhangyang Wang, Zhaowen Wang, Ning Xu, Jiaying Liu, and Zongming Guo. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing
Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, and Amit H. Bermano. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Strokestyles: Stroke-based segmentation and stylization of fonts
Daniel Berio, Frederic Fol Leymarie, Paul Asente, and Jose Echevarria. 2022 · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models. 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
CLIPFont: Text Guided Vector WordArt Generation.. In BMVC . 543
Yiren Song and Yuxuan Zhang. 2022 · 2022
Earlier work this paper cites.
Few-shot font generation by learning fine-grained local styles. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 7895–7904
Licheng Tang, Yiyang Cai, Jiaming Liu, Zhibin Hong, Mingming Gong, Minhu Fan, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. 2022 · 2022
Earlier work this paper cites.
Break-a-scene: Extracting multiple concepts from a single image. In SIGGRAPH Asia 2023 Conference Papers . 1–12
Omri Avrahami, Kfir Aberman, Ohad Fried, Daniel Cohen-Or, and Dani Lischinski. 2023 · 2023
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 18392–18402
Tim Brooks, Aleksander Holynski, and Alexei A Efros. 2023 · 2023
Earlier work this paper cites.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. 2023 · 2023
Earlier work this paper cites.
Diffusion self-guidance for controllable image generation
Dave Epstein, Allan Jabri, Ben Poole, Alexei Efros, and Aleksander Holynski. 2023 · 2023
Earlier work this paper cites.
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. In The Eleventh International Conference on Learning Representations
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. 2023 · 2023
Earlier work this paper cites.
Prompt-to-Prompt Image Editing with Cross-Attention Control. In The Eleventh International Conference on Learning Representations
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. 2023 · 2023
Earlier work this paper cites.
Flow Matching for Generative Modeling. In The Eleventh International Conference on Learning Representations
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. 2023 · 2023
Earlier work this paper cites.
Character-Aware Models Improve Visual Text Rendering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 16270–16297
Rosanne Liu, Dan Garrette, Chitwan Saharia, William Chan, Adam Roberts, Sharan Narang, Irina Blok, Rj Mical, Mohammad Norouzi, and Noah Constant. 2023 · 2023
Earlier work this paper cites.
Intelligent Typography: Artistic Text Style Transfer for Complex Texture and Structure
Wendong Mao, Shuai Yang, Huihong Shi, Jiaying Liu, and Zhongfeng Wang. 2023 · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Clipvg: Text-guided image manipulation using differentiable vector graphics. In Proceedings of the AAAI conference on artificial intelligence , Vol. 37. 2312–2320
Yiren Song, Xuning Shao, Kang Chen, Weidong Zhang, Zhongliang Jing, and Minzhe Li. 2023 · 2023
Cited alongside, same era.
DS-Fusion: Artistic Typography via Discriminated and Stylized Diffusion. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . 374–384
Maham Tanveer, Yizhi Wang, Ali Mahdavi-Amiri, and Hao Zhang. 2023 · 2023
Cited alongside, same era.
Anything to Glyph: Artistic Font Synthesis via Text-to-Image Diffusion Model. In SIGGRAPH Asia 2023 Conference Papers . 1–11
Ominicontrol: Minimal and universal control for diffusion transformer
Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, and Xinchao Wang. 2024 · 2024
Later among the works it cites.
AnyText: Multilingual Visual Text Generation and Editing. In The Twelfth International Conference on Learning Representations
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, and Xuansong Xie. 2024 · 2024
Later among the works it cites.
Diffusion feedback helps clip see better
Wenxuan Wang, Quan Sun, Fan Zhang, Yepeng Tang, Jing Liu, and Xinlong Wang. 2024 · 2024
Later among the works it cites.
Q-ALIGN: teaching LMMs for visual scoring via discrete text-defined levels. In Proceedings of the 41st International Conference on Machine Learning . 54015–54029
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Changshuo Wang, Lei Wu, Xiaole Liu, Xiang Li, Lei Meng, and Xiangxu Meng. 2023 · 2023
Cited alongside, same era.
Omnicontrol: Control any joint at any time for human motion generation
Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang. 2023 · 2023
Cited alongside, same era.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023 · 2023
Cited alongside, same era.
Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3836–3847
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023 · 2023
Cited alongside, same era.
Black Forest Labs - Frontier AI Lab
2024 · 2024
Cited alongside, same era.
Intelligent Artistic Typography: A Comprehensive Review of Artistic Text Design and Generation
Yuhang Bai, Zichuan Huang, Wenshuo Gao, Shuai Yang, Jiaying Liu, et al · 2024
Cited alongside, same era.
PaddleOCR: An open-source optical character recognition (OCR) tool
Baidu. 2024 · 2024
Cited alongside, same era.
FLUX.1 Fill [dev]
black-forest labs. November 1, 2024 · 2024
Cited alongside, same era.
Glyphcontrol: Glyph conditional control for visual text generation
Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu, and Kai Chen. 2024 · 2024
Later among the works it cites.
Ssr-encoder: Encoding selective subject representation for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8069–8078
Yuxuan Zhang, Yiren Song, Jiaming Liu, Rui Wang, Jinpeng Yu, Hao Tang, Huaxia Li, Xu Tang, Yao Hu, Han Pan, et al · 2024
Later among the works it cites.
ColorPeel: Color Prompt Learning with Diffusion Models via Color and Shape Disentanglement. In European Conference on Computer Vision . Springer, 456–472
Muhammad Atif Butt, Kai Wang, Javier Vazquez-Corral, and Joost van de Weijer. 2025 · 2025
Closest in time.
Posta: A go-to framework for customized artistic poster generation
Haoyu Chen, Xiaojie Xu, Wenbo Li, Jingjing Ren, Tian Ye, Songhua Liu, Ying-Cong Chen, Lei Zhu, and Xinchao Wang. 2025 · 2025
Closest in time.
Experiment with Gemini 2.0 Flash native image generation
Google AI. 2025 · 2025
Closest in time.
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
Alec Helbling, Tuna Han Salih Meral, Ben Hoover, Pinar Yanardag, and Duen Horng Chau. 2025 · 2025
Closest in time.
DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics
Yihan Hu, Jianing Peng, Yiheng Lin, Ting Liu, Xiaochao Qu, Luoqi Liu, Yao Zhao, and Yunchao Wei. 2025 · 2025
Closest in time.
Chen Jin, Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, and Philip Teare. 2025 · 2025
Closest in time.
ControlNet + + ++ : Improving Conditional Controls with Efficient Consistency Feedback. In European Conference on Computer Vision . Springer, 129–147
Ming Li, Taojiannan Yang, Huafeng Kuang, Jie Wu, Zhaoning Wang, Xuefeng Xiao, and Chen Chen. 2025 · 2025
Closest in time.
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
Runnan Lu, Yuxuan Zhang, Jiaming Liu, Haofan Wang, and Yiren Song. 2025 · 2025
Closest in time.
Introducing 4O Image Generation
OpenAI. 2025 · 2025
Closest in time.
Generative AI in Fashion: Overview
Wenda Shi, Waikeung Wong, and Xingxing Zou. 2025 · 2025
Closest in time.
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Min Li. 2025b · 2025
Closest in time.
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
Zhendong Wang, Jianmin Bao, Shuyang Gu, Dong Chen, Wengang Zhou, and Houqiang Li. 2025a · 2025
Closest in time.
Less-to-more generalization: Unlocking more controllability by in-context generation
Shaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng, Fei Ding, and Qian He. 2025 · 2025
Closest in time.
ARTIST: Improving the Generation of Text-Rich Images with Disentangled Diffusion Models and Large Language Models. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE, 1268–1278
Jianyi Zhang, Yufan Zhou, Jiuxiang Gu, Curtis Wigington, Tong Yu, Yiran Chen, Tong Sun, and Ruiyi Zhang. 2025b · 2025
Closest in time.
Easycontrol: Adding efficient and flexible control for diffusion transformer
Yuxuan Zhang, Yirui Yuan, Yiren Song, Haofan Wang, and Jiaming Liu. 2025a · 2025
Closest in time.
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
Shitian Zhao, Qilong Wu, Xinyue Li, Bo Zhang, Ming Li, Qi Qin, Dongyang Liu, Kaipeng Zhang, Hongsheng Li, Yu Qiao, et al · 2025
Closest in time.
From Fragment to One Piece: A Survey on AI-Driven Graphic Design
Xingxing Zou, Wen Zhang, and Nanxuan Zhao. 2025 · 2025
Closest in time.