Fetching the paper…
Reading the bibliography…
Poster design is a critical medium for visual communication.
Universal style transfer via feature transforms
Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Adversarial feature matching for text generation
Yizhe Zhang, Zhe Gan, Kai Fan, Zhi Chen, Ricardo Henao, Dinghan Shen, and Lawrence Carin · 2017
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu · 2019
Earlier work this paper cites.
Analyzing and improving the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Agilegan: stylizing portraits by inversion-consistent transfer learning
Guoxian Song, Linjie Luo, Jing Liu, Wan-Chun Ma, Chunpong Lai, Chuanxia Zheng, and Tat-Jen Cham · 2021
Earlier work this paper cites.
Drb-gan: A dynamic resblock generative adversarial network for artistic style transfer
Wenju Xu, Chengjiang Long, Ruisheng Wang, and Guanghui Wang · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Transformer tracking with cyclic shifting window attention
Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang · 2022
Earlier work this paper cites.
Df-gan: A simple and effective baseline for text-to-image synthesis
Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu · 2022
Earlier work this paper cites.
Snow removal in video: A new dataset and a novel method
Haoyu Chen, Jingjing Ren, Jinjin Gu, Hongtao Wu, Xuequan Lu, Haoming Cai, and Lei Zhu · 2023
Earlier work this paper cites.
Wordart designer: User-driven artistic typography synthesis using large language models
Jun-Yan He, Zhi-Qi Cheng, Chenyang Li, Jingdong Sun, Wangmeng Xiang, Xianhui Lin, Xiaoyang Kang, Zengke Jin, Yusen Hu, Bin Luo, et al · 2023
Earlier work this paper cites.
OpenAI. GPT-4V. https://openai.com/index/gpt-4v-system card/, 2023
2023
Earlier work this paper cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Earlier work this paper cites.
Diffbfr: Bootstrapping diffusion model for blind face restoration
Xinmin Qiu, Congying Han, ZiCheng Zhang, Bonan Li, Tiande Guo, and Xuecheng Nie · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Anytext: Multilingual visual text generation and editing
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, and Xuansong Xie · 2023
Cited alongside, same era.
Recraft. Recraft v3. https://www.recraft.ai/about, 2023
2023
Cited alongside, same era.
Anything to glyph: Artistic font synthesis via text-to-image diffusion model
Changshuo Wang, Lei Wu, Xiaole Liu, Xiang Li, Lei Meng, and Xiangxu Meng · 2023
Cited alongside, same era.
Ultrapixel: Advancing ultra-high-resolution image synthesis to new peaks
Jingjing Ren, Wenbo Li, Haoyu Chen, Renjing Pei, Bin Shao, Yong Guo, Long Peng, Fenglong Song, and Lei Zhu · 2024
Later among the works it cites.
Coser: Bridging image and language for cognitive super-resolution
Haoze Sun, Wenbo Li, Jianzhuang Liu, Haoyu Chen, Renjing Pei, Xueyi Zou, Youliang Yan, and Yujiu Yang · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al · 2024
Later among the works it cites.
Ideogram AI. Ideogram v2. https://ideogram.ai/launch, 2024
2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cutmib: Boosting light field super-resolution via multi-view image blending
Zeyu Xiao, Yutong Liu, Ruisheng Gao, and Zhiwei Xiong · 2023
Cited alongside, same era.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Cited alongside, same era.
Intelligent artistic typography: A comprehensive review of artistic text design and generation
Yuhang Bai, Zichuan Huang, Wenshuo Gao, Shuai Yang, Jiaying Liu, et al · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Cited alongside, same era.
Vita: Towards open-source interactive omni multimodal llm
Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Xiong Wang, Di Yin, Long Ma, Xiawu Zheng, et al · 2024
Cited alongside, same era.
Black Forest Labs. FLUX. https://github.com/black-forest labs/flux, 2024
2024
Cited alongside, same era.
Ella: Equip diffusion models with llm for enhanced semantic alignment
Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu · 2024
Cited alongside, same era.
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Later among the works it cites.
General ocr theory: Towards ocr-2.0 via a unified end-to-end model
Haoran Wei, Chenglong Liu, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, et al · 2024
Later among the works it cites.
Image augmentation with controlled diffusion for weakly-supervised semantic segmentation
Wangyu Wu, Tianhong Dai, Xiaowei Huang, Fei Ma, and Jimin Xiao · 2024
Later among the works it cites.
Analyzing large language models’ capability in location prediction
Zhaomin Xiao, Yan Huang, and Eduardo Blanco · 2024
Later among the works it cites.
Short interest trend prediction with large language models
Zhaomin Xiao, Zhelu Mai, Yachen Cui, Zhuoer Xu, and Jiancheng Li · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al · 2024
Later among the works it cites.
mplug-owl3: Towards long image-sequence understanding in multi-modal large language models
Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou · 2024
Later among the works it cites.
Is llama 3 good at identifying emotion? a comprehensive study
Jinran Zhang, Zhelu Mai, Zhuoer Xu, and Zhaomin Xiao · 2024
Later among the works it cites.
Restoreagent: Autonomous image restoration agent via multimodal large language models
Haoyu Chen, Wenbo Li, Jinjin Gu, Jingjing Ren, Sixiang Chen, Tian Ye, Renjing Pei, Kaiwen Zhou, Fenglong Song, and Lei Zhu · 2025
Closest in time.
Triplane-smoothed video dehazing with clip-enhanced generalization
Jingjing Ren, Haoyu Chen, Tian Ye, Hongtao Wu, and Lei Zhu · 2025
Closest in time.
Adaptive patch contrast for weakly supervised semantic segmentation
Wangyu Wu, Tianhong Dai, Zhenhong Chen, Xiaowei Huang, Jimin Xiao, Fei Ma, and Renrong Ouyang · 2025
Closest in time.
Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning
Hanxun Yu, Wentong Li, Song Wang, Junbo Chen, and Jianke Zhu · 2025
Closest in time.