Fetching the paper…
Reading the bibliography…
Generating cognitive-aligned layered SVGs remains challenging due to existing methods' tendencies toward either oversimplified single-layer outputs or optimization-induced shape redundancies.
Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1911–1920
Ajay Jain, Amber Xie, and Pieter Abbeel. 2023 · 1920
Earlier work this paper cites.
Stochastic sampling in computer graphics
Robert L Cook. 1986 · 1986
Earlier work this paper cites.
Convolution surfaces. In Proceedings of the 18th annual conference on Computer graphics and interactive techniques . 251–256
Jules Bloomenthal and Ken Shoemake. 1991 · 1991
Earlier work this paper cites.
Noto Emoji Fonts
Google. 2014 · 2014
Earlier work this paper cites.
A learned representation for scalable vector graphics. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7930–7939
Raphael Gontijo Lopes, David Ha, Douglas Eck, and Jonathon Shlens. 2019 · 2019
Earlier work this paper cites.
Deepsvg: A hierarchical generative network for vector graphics animation
Alexandre Carlier, Martin Danelljan, Alexandre Alahi, and Radu Timofte. 2020 · 2020
Earlier work this paper cites.
Deep vectorization of technical drawings. In European conference on computer vision . Springer, 582–598
Vage Egiazarian, Oleg Voynov, Alexey Artemov, Denis Volkhonskiy, Aleksandr Safin, Maria Taktasheva, Denis Zorin, and Evgeny Burnaev. 2020 · 2020
Earlier work this paper cites.
Differentiable vector graphics rasterization for editing and learning
Tzu-Mao Li, Michal Lukáč, Michaël Gharbi, and Jonathan Ragan-Kelley. 2020 · 2020
Earlier work this paper cites.
Multi-modal attention for speech emotion recognition
Zexu Pan, Zhaojie Luo, Jichen Yang, and Haizhou Li. 2020 · 2020
Earlier work this paper cites.
Clipdraw: Exploring text-to-drawing synthesis through language-image encoders
Kevin Frans, LB Soros, and Olaf Witkowski. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Paint transformer: Feed forward neural painting with stroke prediction. In Proceedings of the IEEE/CVF international conference on computer vision . 6598–6607
Songhua Liu, Tianwei Lin, Dongliang He, Fu Li, Ruifeng Deng, Xin Li, Errui Ding, and Hao Wang. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Im2vec: Synthesizing vector graphics without vector supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7342–7351
Pradyumna Reddy, Michael Gharbi, Michal Lukac, and Niloy J Mitra. 2021 · 2021
Earlier work this paper cites.
Deepvecfont: synthesizing high-quality vector fonts via dual-modality learning
Yizhi Wang and Zhouhui Lian. 2021 · 2021
Earlier work this paper cites.
Towards layer-wise image vectorization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16314–16323
Xu Ma, Yuqian Zhou, Xingqian Xu, Bin Sun, Valerii Filev, Nikita Orlov, Yun Fu, and Humphrey Shi. 2022 · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
Intelli-Paint: Towards developing more human-intelligible painting agents. In European Conference on Computer Vision . Springer, 685–701
Jaskirat Singh, Cameron Smith, Jose Echevarria, and Liang Zheng. 2022 · 2022
Earlier work this paper cites.
Cliptexture: Text-driven texture synthesis. In Proceedings of the 30th ACM International Conference on Multimedia . 5468–5476
Yiren Song. 2022 · 2022
Earlier work this paper cites.
CLIPFont: Text Guided Vector WordArt Generation.. In BMVC . 543
Yiren Song and Yuxuan Zhang. 2022 · 2022
Earlier work this paper cites.
Clipasso: Semantically-aware object sketching
Yael Vinker, Ehsan Pajouheshgar, Jessica Y Bo, Roman Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir. 2022 · 2022
Cited alongside, same era.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al · 2023
Cited alongside, same era.
Image vectorization and editing via linear gradient layer decomposition
Zheng-Jun Du, Liang-Fu Kang, Jianchao Tan, Yotam Gingold, and Kun Xu. 2023 · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2023 · 2023
Cited alongside, same era.
Stroke-based Neural Painting and Stylization with Dynamically Predicted Painting Region. In Proceedings of the 31st ACM International Conference on Multimedia . 7470–7480
NIVeL: Neural Implicit Vector Layers for Text-to-Vector Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4589–4597
Vikas Thamizharasan, Difan Liu, Matthew Fisher, Nanxuan Zhao, Evangelos Kalogerakis, and Michal Lukac. 2024 · 2024
Later among the works it cites.
GRID: Visual Layout Generation
Cong Wan, Xiangyang Luo, Zijian Cai, Yiren Song, Yunlong Zhao, Yifan Bai, Yuhang He, and Yihong Gong. 2024 · 2024
Later among the works it cites.
Instantid: Zero-shot identity-preserving generation in seconds
Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and Anthony Chen. 2024a · 2024
Later among the works it cites.
Layered Image Vectorization via Semantic Simplification
Zhenyu Wang, Jianxi Huang, Zhida Sun, Daniel Cohen-Or, and Min Lu. 2024b · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Teng Hu, Ran Yi, Haokun Zhu, Liang Liu, Jinlong Peng, Yabiao Wang, Chengjie Wang, and Lizhuang Ma. 2023 · 2023
Cited alongside, same era.
Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4195–4205
William Peebles and Saining Xie. 2023 · 2023
Cited alongside, same era.
Starvector: Generating scalable vector graphics code from images
Juan A Rodriguez, Shubham Agarwal, Issam H Laradji, Pau Rodriguez, David Vazquez, Christopher Pal, and Marco Pedersoli. 2023 · 2023
Cited alongside, same era.
Clipvg: Text-guided image manipulation using differentiable vector graphics. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 2312–2320
Yiren Song, Xuning Shao, Kang Chen, Weidong Zhang, Zhongliang Jing, and Minzhe Li. 2023 · 2023
Cited alongside, same era.
IconShop: Text-Based Vector Icon Synthesis with Autoregressive Transformers
Ronghuan Wu, Wanchao Su, Kede Ma, and Jing Liao. 2023 · 2023
Cited alongside, same era.
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. 2023 · 2023
Cited alongside, same era.
Diffsketcher: Text guided vector sketch synthesis through latent diffusion models
Ximing Xing, Chuang Wang, Haitao Zhou, Jing Zhang, Qian Yu, and Dong Xu. 2023 · 2023
Cited alongside, same era.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023 · 2023
Cited alongside, same era.
Ximing Xing, Juncheng Hu, Guotao Liang, Jing Zhang, Dong Xu, and Qian Yu. 2024a · 2024
Later among the works it cites.
Empowering LLMs to Understand and Generate Complex Vector Graphics
Ximing Xing, juncheng Hu, Liang Zhang, Jing Guotao, Dong Xu, and Qian Yu. 2024b · 2024
Later among the works it cites.
Text-to-vector generation with neural path representation
Peiying Zhang, Nanxuan Zhao, and Jing Liao. 2024e · 2024
Later among the works it cites.
Fast Personalized Text to Image Synthesis with Attention Injection. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 6195–6199
Yuxuan Zhang, Yiren Song, Jinpeng Yu, Han Pan, and Zhongliang Jing. 2024b · 2024
Later among the works it cites.
Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model
Yuxuan Zhang, Lifu Wei, Qing Zhang, Yiren Song, Jiaming Liu, Huaxia Li, Xu Tang, Yao Hu, and Haibo Zhao. 2024c · 2024
Later among the works it cites.
Stable-Hair: Real-World Hair Transfer via Diffusion Model
Yuxuan Zhang, Qing Zhang, Yiren Song, and Jiaming Liu. 2024d · 2024
Later among the works it cites.
Transanimate: Taming layer diffusion to generate rgba video
Xuewei Chen, Zhimin Chen, and Yiren Song. 2025 · 2025
Closest in time.
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
Yan Gong, Yiren Song, Yicheng Li, Chenglin Li, and Yin Zhang. 2025 · 2025
Closest in time.
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
Hailong Guo, Bohan Zeng, Yiren Song, Wentao Zhang, Chuang Zhang, and Jiaming Liu. 2025 · 2025
Closest in time.
Photodoodle: Learning artistic image editing from few-shot pairwise data
Shijie Huang, Yiren Song, Yuxuan Zhang, Hailong Guo, Xueyin Wang, Mike Zheng Shou, and Jiaming Liu. 2025 · 2025
Closest in time.
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
Runnan Lu, Yuxuan Zhang, Jiaming Liu, Haofan Wang, and Yiren Song. 2025 · 2025
Closest in time.
WordCon: Word-level Typography Control in Scene Text Rendering
Wenda Shi, Yiren Song, Zihan Rao, Dengming Zhang, Jiaming Liu, and Xingxing Zou. 2025 · 2025
Closest in time.
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
Yiren Song, Cheng Liu, and Mike Zheng Shou. 2025a · 2025
Closest in time.
Omniconsistency: Learning style-agnostic consistency from paired stylization data
Yiren Song, Cheng Liu, and Mike Zheng Shou. 2025b · 2025
Closest in time.
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
Zitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu, Xiangtai Li, and Yiren Song. 2025 · 2025
Closest in time.
Easycontrol: Adding efficient and flexible control for diffusion transformer
Yuxuan Zhang, Yirui Yuan, Yiren Song, Haofan Wang, and Jiaming Liu. 2025 · 2025
Closest in time.