Fetching the paper…
Reading the bibliography…
Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently.
Visualizing Data using t-SNE
Laurens van der Maaten and Geoffrey E. Hinton · 2008
Earlier work this paper cites.
TranslatAR: A Mobile Augmented Reality Translator
Victor Fragoso, Steffen Gauglitz, Shane Zamora, Jim Kleban, and Matthew A. Turk · 2011
Earlier work this paper cites.
ICDAR 2013 Robust Reading Competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, M. Iwamura, Lluís Gómez i Bigorda, Sergi Robles Mestre, Joan Mas Romeu, David Fernández Mota, Jon Almazán, and Lluís-Pere de las Heras · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-scale Image Recognition
K Simonyan and A Zisserman · 2015
Earlier work this paper cites.
Synthetic Data for Text Localisation in Natural Images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Earlier work this paper cites.
Pyramid Scene Parsing Network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia · 2016
Earlier work this paper cites.
V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation
Fausto Milletarì, Nassir Navab, and Seyed-Ahmad Ahmadi · 2016
Earlier work this paper cites.
Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization
Xun Huang and Serge J. Belongie · 2017
Earlier work this paper cites.
ICDAR2017 Robust Reading Challenge on Multi-Lingual Scene Text Detection and Script Identification - RRC-MLT
Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon, Wafa Khlif, Muhammad Muzzamil Luqman, Jean-Christophe Burie, Cheng-Lin Liu, and Jean-Marc Ogier · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
CBAM: Convolutional Block Attention Module
Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In-So Kweon · 2018
Earlier work this paper cites.
Editing Text in the Wild
Liang Wu, Chengquan Zhang, Jiaming Liu, Junyu Han, Jingtuo Liu, Errui Ding, and Xiang Bai · 2019
Earlier work this paper cites.
Analyzing and Improving the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2019
Earlier work this paper cites.
What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
Jeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park, Dongyoon Han, Sangdoo Yun, Seong Joon Oh, and Hwalsuk Lee · 2019
Earlier work this paper cites.
SEED: Semantics Enhanced Encoder-decoder Framework for Scene Text Recognition
Zhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou, and Weiping Wang · 2020
Earlier work this paper cites.
SwapText: Image Based Texts Transfer in Scenes
Qiangpeng Yang, Hongsheng Jin, Jun Huang, and Wei Lin · 2020
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Cited alongside, same era.
TextStyleBrush: Transfer of Text Aesthetics From a Single Example
Praveen Krishnan, Rama Kovvuri, Guan Pang, Boris Vassilev, and Tal Hassner · 2021
Cited alongside, same era.
PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text Recognition
Zhi Qiao, Yu Zhou, Jin Wei, Wei Wang, Yuan Zhang, Ning Jiang, Hongbin Wang, and Weiping Wang · 2021
Cited alongside, same era.
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal and Alex Nichol · 2021
Cited alongside, same era.
STRIVE: Scene Text Replacement In Videos
Jeyasri Subramanian, Varnith Chordia, Eugene Bart, Shaobo Fang, Kelly Guan, Raja Bala, et al · 2021
Exploring Stroke-Level Modifications for Scene Text Editing
Yadong Qu, Qingfeng Tan, Hongtao Xie, Jianjun Xu, Yuxin Wang, and Yongdong Zhang · 2023
Later among the works it cites.
Letter Embedding Guidance Diffusion Model for Scene Text Editing
Changshuo Wang, L. Wu, Xu Chen, Xiang Li, Lei Meng, and Xiangxu Meng · 2023
Later among the works it cites.
Improving Diffusion Models for Scene Text Editing with Dual Encoders
Jiabao Ji, Guanhua Zhang, Zhaowen Wang, Bairu Hou, Zhifei Zhang, Brian L Price, and Shiyu Chang · 2023
Later among the works it cites.
Null-text Inversion for Editing Real Images using Guided Diffusion Models
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or · 2023
Later among the works it cites.
Yiming Zhao and Zhouhui Lian · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition
Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang · 2021
Cited alongside, same era.
Vision Transformer for Fast and Efficient Scene Text Recognition
Rowel Atienza · 2021
Cited alongside, same era.
TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text Representation
Wei Wang, Yu Zhou, Jiahao Lv, Dayan Wu, Guoqing Zhao, Ning Jiang, and Weipinng Wang · 2022
Cited alongside, same era.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Later among the works it cites.
Imagic: Text-based real image editing with diffusion models
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani · 2023
Later among the works it cites.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng · 2023
Later among the works it cites.
Cross-Image Attention for Zero-Shot Appearance Transfer
Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor, and Daniel Cohen-Or · 2023
Later among the works it cites.
ICDAR 2023 Competition on Hierarchical Text Detection and Recognition
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis · 2023
Later among the works it cites.
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
Yan Shu, Weichao Zeng, Zhenhang Li, Fangmin Zhao, and Yu Zhou · 2024
Closest in time.
On Manipulating Scene Text in the Wild with Diffusion Models
Joshua Santoso, Christian Simon, et al · 2024
Closest in time.
TextDiffuser: Diffusion Models as Text Painters
Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei · 2024
Closest in time.
AnyText: Multilingual Visual Text Generation And Editing
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, and Xuansong Xie · 2024
Closest in time.
DiffUTE: Universal Text Editing Diffusion Model
Haoxing Chen, Zhuoer Xu, Zhangxuan Gu, Yaohui Li, Changhua Meng, Huijia Zhu, Weiqiang Wang, et al · 2024
Closest in time.
GlyphControl: Glyph Conditional Control for Visual Text Generation
Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu, and Kai Chen · 2024
Closest in time.
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
Zhenhang Li, Yan Shu, Weichao Zeng, Dongbao Yang, and Yu Zhou · 2024
Closest in time.
MotionEditor: Editing Video Motion via Content-Aware Diffusion
Shuyuan Tu, Qi Dai, Zhi-Qi Cheng, Han Hu, Xintong Han, Zuxuan Wu, and Yu-Gang Jiang · 2024
Closest in time.
Diff-Font: Diffusion Model for Robust One-shot Font Generation
Haibin He, Xinyuan Chen, Chaoyue Wang, Juhua Liu, Bo Du, Dacheng Tao, and Qiao Yu · 2024
Closest in time.