Fetching the paper…
Reading the bibliography…
Visual text is a crucial component in both document and scene images, conveying rich semantic information and attracting significant attention in the computer vision community.
G. Jiang, S. Wang, T. Ge, Y. Jiang, Y. Wei, and D. Lian, “Self-supervised text erasing with controllable image synthesis,” in ACM MM , 2022, pp. 1973–1983
1983
Earlier work this paper cites.
Z. Wang, E. Simoncelli, and A. Bovik, “Multiscale structural similarity for image quality assessment,” in ACSSC , vol. 2, 2003, pp. 1398–1402 Vol.2
2003
Earlier work this paper cites.
N. Ezaki, M. Bulacu, and L. Schomaker, “Text detection from natural scene images: Towards a system for visually impaired persons,” in ICPR , vol. 2. IEEE, 2004, pp. 683–686
2004
Earlier work this paper cites.
Y.-C. Tsoi and M. S. Brown, “Multi-view document rectification using boundary,” in CVPR , 2007, pp. 1–8
2007
Earlier work this paper cites.
L. Zhang, A. M. Yip, M. S. Brown, and C. L. Tan, “A unified framework for document restoration using inpainting and shape-from-shading,” PR , vol. 42, pp. 2961–2978, 2009
2009
Earlier work this paper cites.
V. Fragoso, S. Gauglitz, S. Zamora, J. Kleban, and M. Turk, “TranslatAR: A mobile augmented reality translator,” in WACV , 2011, pp. 497–502
2011
Earlier work this paper cites.
G. Meng, C. Pan, S. Xiang, J. Duan, and N. Zheng, “Metric rectification of curved document images,” TPAMI , vol. 34, no. 4, pp. 707–722, 2011
2011
Earlier work this paper cites.
Q. Ye and D. Doermann, “Text detection and recognition in imagery: A survey,” TPAMI , vol. 37, no. 7, pp. 1480–1500, 2014
2014
Earlier work this paper cites.
K. Inai, M. Pålsson, V. Frinken, Y. Feng, and S. Uchida, “Selective concealment of characters for privacy protection,” in CVPR , 2014, pp. 333–338
2014
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman, “Synthetic data and artificial neural networks for natural scene text recognition,” in NIPSW , 2014
2014
Earlier work this paper cites.
C. Peyrard, M. Baccouche, F. Mamalet, and C. Garcia, “ICDAR2015 competition on text image super-resolution,” ICDAR , pp. 1201–1205, 2015
2015
Earlier work this paper cites.
M. Hradiš, J. Kotera, P. Zemcık, and F. Šroubek, “Convolutional neural networks for direct text deblurring,” in BMVC , vol. 10, no. 2, 2015
2015
Earlier work this paper cites.
Y. Zhu, C. Yao, and X. Bai, “Scene text detection and recognition: Recent advances and future trends,” Front. Comput. Sci. , vol. 10, pp. 19–36, 2016
2016
Earlier work this paper cites.
X.-C. Yin, Z.-Y. Zuo, S. Tian, and C.-L. Liu, “Text detection, tracking and recognition in video: A comprehensive survey,” TIP , vol. 25, no. 6, pp. 2752–2773, 2016
2016
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” TPAMI , vol. 39, no. 11, pp. 2298–2304, 2016
2016
Earlier work this paper cites.
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. P. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi, “Photo-realistic single image super-resolution using a generative adversarial network,” CVPR , pp. 105–114, 2016
2016
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” TPAMI , vol. 39, no. 11, pp. 2298–2304, 2016
2016
Earlier work this paper cites.
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in CVPR , 2016, pp. 2315–2324
2016
Earlier work this paper cites.
S. Das, G. Mishra, A. Sudharshana, and R. Shilkrot, “The common fold: Utilizing the four-fold to dewarp printed documents from a single image,” ACM SDE , 2017
2017
Earlier work this paper cites.
T. Nakamura, A. Zhu, K. Yanai, and S. Uchida, “Scene text eraser,” in ICDAR , vol. 1, 2017, pp. 832–837
2017
Earlier work this paper cites.
Y. Jiang, Z. Lian, Y. Tang, and J. Xiao, “DCfont: an end-to-end deep chinese font generation system,” in SIGGRAPH Asia , 2017, pp. 1–4
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local nash equilibrium,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in CVPR , 2017, pp. 1125–1134
2017
Earlier work this paper cites.
H. Abu Alhaija, S. K. Mustikovela, L. Mescheder, A. Geiger, and C. Rother, “Augmented reality meets computer vision: Efficient data generation for urban driving scenes,” IJCV , vol. 126, pp. 961–972, 2018
2018
Earlier work this paper cites.
S. Qin, P. Ren, S. Kim, and R. Manduchi, “Robust and accurate text stroke segmentation,” in WACV , 2018, pp. 242–250
2018
Earlier work this paper cites.
S. Qin, J. Wei, and R. Manduchi, “Automatic semantic content removal by learning to neglect,” in BMVC , 2018
2018
Earlier work this paper cites.
Y. Zhang, Y. Zhang, and W. Cai, “Separating style and content for generalized style transfer,” in CVPR , 2018, pp. 8447–8455
2018
Earlier work this paper cites.
F. Zhan, S. Lu, and C. Xue, “Verisimilar image synthesis for accurate detection and recognition of texts in scenes,” in ECCV , 2018, pp. 249–266
2018
Earlier work this paper cites.
K. Ma, Z. Shu, X. Bai, J. Wang, and D. Samaras, “DocUNet: Document image unwarping via a stacked U-Net,” CVPR , pp. 4700–4709, 2018
2018
Earlier work this paper cites.
N. Kligler, S. Katz, and A. Tal, “Document enhancement using visibility detection,” CVPR , pp. 2374–2382, 2018
2018
Earlier work this paper cites.
I. Pratikakis, K. Zagori, P. Kaddas, and B. Gatos, “ICFHR 2018 competition on handwritten document image binarization (H-DIBCO 2018),” in ICFHR . IEEE, 2018, pp. 489–493
2018
Earlier work this paper cites.
X. Liu, G. Meng, and C. Pan, “Scene text detection and recognition with advances in deep learning: A survey,” IJDAR , vol. 22, pp. 143–162, 2019
2019
Earlier work this paper cites.
R. Nakao, B. K. Iwana, and S. Uchida, “Selective super-resolution for scene text images,” ICDAR , pp. 401–406, 2019
2019
Earlier work this paper cites.
S. He and L. Schomaker, “DeepOtsu: Document enhancement and binarization using iterative deep learning,” PR , vol. 91, pp. 379–390, 2019
2019
Earlier work this paper cites.
W. R. Huang, Y. Qi, Q. Li, J. Degange, and Y. Llp, “DeepErase: Weakly supervised ink artifact removal in document text images,” WACV , pp. 3511–3519, 2019
2019
Earlier work this paper cites.
O. Tursun, R. Zeng, S. Denman, S. Sivapalan, S. Sridharan, and C. Fookes, “MTRNet: A generic scene text eraser,” in ICDAR , 2019, pp. 39–44
2019
Earlier work this paper cites.
T. N. Nakamura, A. Zhu, and S. Uchida, “Scene text magnifier,” in ICDAR , 2019, pp. 825–830
2019
Earlier work this paper cites.
R. Gomez, A. F. Biten, L. Gomez, J. Gibert, D. Karatzas, and M. Rusiñol, “Selective style transfer for text,” in ICDAR , 2019, pp. 805–812
2019
Earlier work this paper cites.
L. Wu, C. Zhang, J. Liu, J. Han, J. Liu, E. Ding, and X. Bai, “Editing text in the wild,” in ACM MM , 2019, pp. 1500–1508
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Wang, F. Su, and Y. Qian, “Text-attentional conditional generative adversarial network for super-resolution of text images,” ICME , pp. 1024–1029, 2019
2019
Earlier work this paper cites.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of StyleGAN,” CVPR , pp. 8107–8116, 2019
2019
Earlier work this paper cites.
S. Das, K. Ma, Z. Shu, D. Samaras, and R. Shilkrot, “DewarpNet: Single-image document unwarping with stacked 3D and 2D regression networks,” ICCV , pp. 131–140, 2019
2019
Earlier work this paper cites.
S. Zhang, Y. Liu, L. Jin, Y. Huang, and S. Lai, “EnsNet: Ensconce text in the wild,” in AAAI , vol. 33, no. 01, 2019, pp. 801–808
2019
Earlier work this paper cites.
F. Zhan, H. Zhu, and S. Lu, “Spatial fusion GAN for image synthesis,” in CVPR , 2019, pp. 3653–3662
2019
Earlier work this paper cites.
S. Fang, H. Xie, J. Chen, J. Tan, and Y. Zhang, “Learning to draw text in natural images with conditional adversarial networks.” in IJCAI , 2019, pp. 715–722
2019
Earlier work this paper cites.
H. Lin, P. Yang, and F. Zhang, “Review of scene text detection and recognition,” Arch. Comput. Methods Eng. , vol. 27, no. 2, pp. 433–454, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
Z. Qiao, Y. Zhou, D. Yang, Y. Zhou, and W. Wang, “SEED: Semantics enhanced encoder-decoder framework for scene text recognition,” in CVPR , 2020, pp. 13 528–13 537
2020
Earlier work this paper cites.
Y.-H. Lin, W.-C. Chen, and Y.-Y. Chuang, “BEDSR-Net: A deep shadow removal network from a single document image,” CVPR , pp. 12 902–12 911, 2020
2020
Earlier work this paper cites.
M. A. Souibgui and Y. Kessentini, “DE-GAN: A conditional generative adversarial network for document enhancement,” TPAMI , vol. 44, no. 3, pp. 1180–1191, 2020
2020
Earlier work this paper cites.
O. Tursun, S. Denman, R. Zeng, S. Sivapalan, S. Sridharan, and C. Fookes, “MTRNet++: One-stage mask-based scene text eraser,” CVIU , vol. 201, p. 103066, 2020
2020
Earlier work this paper cites.
Q. Yang, J. Huang, and W. Lin, “SwapText: Image based texts transfer in scenes,” in CVPR , 2020, pp. 14 700–14 709
2020
Earlier work this paper cites.
S. Bonechi, M. Bianchini, F. Scarselli, and P. Andreini, “Weak supervision for generating pixel–level annotations in scene text segmentation,” PRL , vol. 138, pp. 1–7, 2020
2020
Earlier work this paper cites.
Y. Mou, L. Tan, H. Yang, J. Chen, L. Liu, R. Yan, and Y. Huang, “PlugNet: Degradation aware scene text recognition supervised by a pluggable super-resolution unit,” in ECCV , 2020
2020
Earlier work this paper cites.
W. Wang, E. Xie, X. Liu, W. Wang, D. Liang, C. Shen, and X. Bai, “Scene text image super-resolution in the wild,” in ECCV . Springer, 2020, pp. 650–666
2020
Earlier work this paper cites.
G.-W. Xie, F. Yin, X.-Y. Zhang, and C.-L. Liu, “Dewarping document image by displacement flow estimation with fully convolutional network,” in DAS , 2020
2020
Earlier work this paper cites.
C. Liu, Y. Liu, L. Jin, S. Zhang, C. Luo, and Y. Wang, “EraseNet: End-to-end text removal in the wild,” TIP , vol. 29, pp. 8760–8775, 2020
2020
Earlier work this paper cites.
P. Roy, S. Bhattacharya, S. Ghosh, and U. Pal, “STEFANN: Scene text editor using font adaptive neural network,” in CVPR , 2020, pp. 13 228–13 237
2020
Earlier work this paper cites.
M. Liao, B. Song, S. Long, M. He, C. Yao, and X. Bai, “SynthText3D: Synthesizing scene text images from 3D virtual worlds,” Sci. China Inf. Sci. , vol. 63, pp. 1–14, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Fogel, H. Averbuch-Elor, S. Cohen, S. Mazor, and R. Litman, “ScrabbleGAN: Semi-supervised varying length handwritten text generation,” in CVPR , 2020, pp. 4324–4333
2020
Earlier work this paper cites.
B. Wang and C. P. Chen, “Local water-filling algorithm for shadow detection and removal of document images,” Sensors , vol. 20, no. 23, p. 6929, 2020
2020
Earlier work this paper cites.
X. Liu, G. Meng, B. Fan, S. Xiang, and C. Pan, “Geometric rectification of document images using adversarial gated unwarping network,” PR , vol. 108, p. 107576, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Khan, R. Sarkar, and A. F. Mollah, “Deep learning approaches to scene text detection: A comprehensive review,” ARTIF INTELL REV , vol. 54, pp. 3239–3298, 2021
2021
Earlier work this paper cites.
X. Chen, L. Jin, Y. Zhu, C. Luo, and T. Wang, “Text recognition in the wild: A survey,” CSUR , vol. 54, no. 2, pp. 1–35, 2021
2021
Earlier work this paper cites.
S. Long, X. He, and C. Yao, “Scene text detection and recognition: The deep learning era,” IJCV , vol. 129, no. 1, pp. 161–184, 2021
2021
Earlier work this paper cites.
S. Dey and P. Jawanpuria, “Light-weight document image cleanup using perceptual loss,” in ICDAR , 2021
2021
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, X. Liu, D. Liang, Z. Yang, T. Lu, and C. Shen, “PAN++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text,” TPAMI , vol. 44, no. 9, pp. 5349–5367, 2021
2021
Earlier work this paper cites.
H. Feng, Y. Wang, W. gang Zhou, J. Deng, and H. Li, “DocTr: Document image transformer for geometric unwarping and illumination correction,” ACM MM , 2021
2021
Earlier work this paper cites.
C. Wang, S. Zhao, L. Zhu, K. Luo, Y. Guo, J. Wang, and S. Liu, “Semi-supervised pixel-level scene text segmentation by mutually guided network,” TIP , vol. 30, pp. 8212–8221, 2021
2021
Cited alongside, same era.
X. Xu, Z. Zhang, Z. Wang, B. Price, Z. Wang, and H. Shi, “Rethinking text segmentation: A novel dataset and a text-specific refinement approach,” in CVPR , 2021, pp. 12 045–12 055
2021
Cited alongside, same era.
J. Ma, S. Guo, and L. Zhang, “Text prior guided scene text image super-resolution,” TIP , vol. 32, pp. 1341–1353, 2021
2021
Cited alongside, same era.
J. Chen, B. Li, and X. Xue, “Scene text telescope: Text-focused scene image super-resolution,” CVPR , pp. 12 021–12 030, 2021
2021
Cited alongside, same era.
C. Zhao, S. Feng, B. N. Zhao, Z. Ding, J. Wu, F. Shen, and H. T. Shen, “Scene text image super-resolution via parallelly contextual attention network,” ACM MM , 2021
H. Chen, Z. Xu, Z. Gu, Y. Li, C. Meng, H. Zhu, W. Wang et al. , “Diffute: Universal text editing diffusion model,” NeurIPS , vol. 36, pp. 63 062–63 074, 2023
2023
Later among the works it cites.
Y. Tuo, W. Xiang, J.-Y. He, Y. Geng, and X. Xie, “Anytext: Multilingual visual text generation and editing,” in ICLR , 2023
2023
Later among the works it cites.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in ICCV , 2023, pp. 3836–3847
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Z. Tang, T. Miyazaki, Y. Sugaya, and S. Omachi, “Stroke-based scene text erasing using synthetic data for training,” TIP , vol. 30, pp. 9306–9320, 2021
2021
Cited alongside, same era.
P. Keserwani and P. P. Roy, “Text region conditional generative adversarial network for text concealment in the wild,” TCSVT , vol. 32, no. 5, pp. 3152–3163, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
L. Zhao, C. Chen, and J. Huang, “Deep learning-based forgery attack on document images,” TIP , vol. 30, pp. 7964–7979, 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
L. Zhang, X. Chen, Y. Xie, and Y. Lu, “Scene text transfer for cross-language,” in ICIG , 2021, pp. 552–564
2021
Cited alongside, same era.
J. Lee, Y. Kim, S. Kim, M. Yim, S. Shin, G. Lee, and S. Park, “RewriteNet: Reliable scene text editing with implicit decomposition of text contents and styles,” CVPRW , 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
R. Liu, D. Garrette, C. Saharia, W. Chan, A. Roberts, S. Narang, I. Blok, R. Mical, M. Norouzi, and N. Constant, “Character-aware models improve visual text rendering,” in ACL , 2023
2023
Later among the works it cites.
F. Zhan, Y. Yu, R. Wu, J. Zhang, S. Lu, L. Liu, A. Kortylewski, C. Theobalt, and E. Xing, “Multimodal image synthesis and editing: The generative AI era,” TPAMI , vol. 45, no. 12, pp. 15 098–15 119, 2023
2023
Later among the works it cites.
J. Ma, Z. Liang, W. Xiang, X. Yang, and L. Zhang, “A benchmark for Chinese-English scene text image super-resolution,” in ICCV , 2023, pp. 19 452–19 461
2023
Later among the works it cites.
L. Zhang, Y. He, Q. Zhang, Z. Liu, X. Zhang, and C. Xiao, “Document image shadow removal guided by color-aware background,” in CVPR , 2023, pp. 1818–1827
2023
Later among the works it cites.
G. Lyu, K. Liu, A. Zhu, S. Uchida, and B. K. Iwana, “FETNet: Feature erasing and transferring network for scene text removal,” PR , vol. 140, p. 109531, 2023
2023
Later among the works it cites.
Y. Yang, D. Gui, Y. Yuan, W. Liang, H. Ding, H. Hu, and K. Chen, “GlyphControl: Glyph conditional control for visual text generation,” in NeurIPS , 2023, pp. 44 050–44 066
2023
Later among the works it cites.
C. Huang, X. Peng, D. Liu, and Y. Lu, “Text image super-resolution guided by text structure and embedding priors,” TOMM , vol. 19, pp. 1 – 18, 2023
2023
Later among the works it cites.
X. Du, Z. Zhou, Y. Zheng, T. Ma, X. Wu, and C. Jin, “Modeling stroke mask for end-to-end text erasing,” in WACV , 2023, pp. 6151–6159
2023
Later among the works it cites.
Y. Wang, H. Xie, Z. Wang, Y. Qu, and Y. Zhang, “What is the real need for scene text removal? Exploring the background integrity and erasure exhaustivity properties,” TIP , 2023
2023
Later among the works it cites.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in ICCV , 2023, pp. 3836–3847
2023
Later among the works it cites.
Y. Li, H. Wang, Q. Jin, J. Hu, P. Chemerys, Y. Fu, Y. Wang, S. Tulyakov, and J. Ren, “SnapFusion: Text-to-image diffusion model on mobile devices within two seconds,” in NeurIPS , 2023, pp. 20 662–20 678
2023
Later among the works it cites.
2024
Later among the works it cites.
H. Wang, M. Liao, Z. Xie, W. Liu, and X. Bai, “Partial scene text retrieval,” TPAMI , 2024
2024
Later among the works it cites.
J. Zhang, D. Peng, C. Liu, P. Zhang, and L. Jin, “DocRes: A generalist model toward unifying document image restoration tasks,” in CVPR , 2024, pp. 15 654–15 664
2024
Later among the works it cites.
T. Wu, K. Ma, J. Liang, Y. Yang, and L. Zhang, “A comprehensive study of multimodal large language models for image quality assessment,” in ECCV . Springer, 2024, pp. 143–160
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Kong, W. Ma, L. Jin, and Y. Xue, “Garden: Generative prior guided network for scene text image super-resolution,” in ICDAR . Springer, 2024, pp. 196–214
2024
Later among the works it cites.
Z. Zhu, L. Zhang, Y. Bai, Y. Wang, and P. Li, “Scene text image super-resolution through multi-scale interaction of structural and semantic priors,” TAI , 2024
2024
Later among the works it cites.
W. Yu, Y. Liu, X. Zhu, H. Cao, X. Sun, and X. Bai, “Turning a CLIP model into a scene text spotter,” TPAMI , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Wang, Y. Xue, and L. Jin, “DocNLC: A document image enhancement framework with normalized and latent contrastive representation for multiple degradations,” in AAAI , vol. 38, no. 6, 2024, pp. 5563–5571
2024
Later among the works it cites.
Z. Yang, B. Liu, Y. Xiong, and G. Wu, “GDB: Gated convolutions-based document binarization,” PR , vol. 146, p. 109989, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
K. Nikolaidou, G. Retsinas, G. Sfikas, and M. Liwicki, “DiffusionPen: Towards controlling the style of handwritten text generation,” in ECCV . Springer, 2024, pp. 417–434
2024
Later among the works it cites.
Z. Yang, D. Peng, Y. Kong, Y. Zhang, C. Yao, and L. Jin, “FontDiffuser: One-shot font generation via denoising diffusion with multi-scale content aggregation and style contrastive learning,” in AAAI , vol. 38, no. 7, 2024, pp. 6603–6611
2024
Later among the works it cites.
Y. Liu and Z. Lian, “QT-Font: High-efficiency font synthesis via quadtree-based diffusion models,” in SIGGRAPH , 2024, pp. 1–11
2024
Later among the works it cites.
G. Yao, K. Zhao, C. Deng, N. Ding, T. Zhao, Y. Tao, and L. Peng, “Geometric-aware control in diffusion model for handwritten chinese font generation,” in ICDAR . Springer, 2024, pp. 3–17
2024
Later among the works it cites.
M. Ye, J. Zhang, J. Liu, C. Liu, B. Yin, C. Liu, B. Du, and D. Tao, “Hi-SAM: Marrying segment anything model for hierarchical text segmentation,” TPAMI , 2024
2024
Later among the works it cites.
C. Noguchi, S. Fukuda, and M. Yamanaka, “Scene text image super-resolution based on text-conditional diffusion models,” in WACV , 2024, pp. 1485–1495
2024
Later among the works it cites.
Z. Zhao, H. Xue, P. Fang, and S. Zhu, “PEAN: A diffusion-based prior-enhanced attention network for scene text image super-resolution,” in ACM MM , 2024, pp. 9769–9778
2024
Later among the works it cites.
Y. Zhang, J. Zhang, H. Li, Z. Wang, L. Hou, D. Zou, and L. Bian, “Diffusion-based blind text image super-resolution,” in CVPR , 2024, pp. 25 827–25 836
2024
Later among the works it cites.
H. Tang, J. Guo, T. Wang, Y. Yu, and C. Wang, “Efficient joint rectification of photometric and geometric distortions in document images,” in ICASSP . IEEE, 2024, pp. 3690–3694
2024
Later among the works it cites.
W. Zeng, Y. Shu, Z. Li, D. Yang, and Y. Zhou, “Textctrl: Diffusion-based scene text editing with prior guidance control,” NeurIPS , 2024
2024
Later among the works it cites.
J. Santoso, C. Simon et al. , “On manipulating scene text in the wild with diffusion models,” in WACV , 2024, pp. 5202–5211
2024
Later among the works it cites.
Y. Zhao and Z. Lian, “Udifftext: A unified framework for high-quality text synthesis in arbitrary images via character-aware diffusion models,” in ECCV . Springer, 2024, pp. 217–233
2024
Later among the works it cites.
B. Zhang, Z. Gao, Y. Qu, and H. Xie, “How control information influences multilingual text image generation and editing?” in NeurIPS , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Zhang, H. Xie, Z. Gao, and Y. Wang, “Choose what you need: Disentangled representation learning for scene text recognition removal and editing,” in CVPR , 2024, pp. 28 358–28 368
2024
Later among the works it cites.
M. Xia, Y. Zhou, R. Yi, Y.-J. Liu, and W. Wang, “A diffusion model translator for efficient image-to-image translation,” TPAMI , 2024
2024
Later among the works it cites.
Z. Liu, W. Liang, Z. Liang, C. Luo, J. Li, G. Huang, and Y. Yuan, “Glyph-ByT5: A customized text encoder for accurate visual text rendering,” in ECCV . Springer, 2024, pp. 361–377
2024
Later among the works it cites.
J. Chen, Y. Huang, T. Lv, L. Cui, Q. Chen, and F. Wei, “TextDiffuser-2: Unleashing the power of language models for text rendering,” in ECCV . Springer, 2024, pp. 386–402
2024
Later among the works it cites.
2024
Later among the works it cites.
F. Yu, Y. Xie, L. Wu, Y. Wen, G. Wang, S. Ren, X. Chen, J. Mao, and W. Li, “DocReal: Robust document dewarping of real-life images via attention-enhanced control point prediction,” in WACV , 2024, pp. 665–674
2024
Later among the works it cites.
Z. Li, Y. Shu, W. Zeng, D. Yang, and Y. Zhou, “First creating backgrounds then rendering texts: A new paradigm for visual text blending,” in ECAI , 2024
2024
Later among the works it cites.
L. Zhang, X. Chen, Y. Wang, Y. Lu, and Y. Qiao, “Brush your text: Synthesize any scene text on images via diffusion model,” in AAAI , vol. 38, no. 7, 2024, pp. 7215–7223
2024
Later among the works it cites.
Y. Zhu, J. Liu, F. Gao, W. Liu, X. Wang, P. Wang, F. Huang, C. Yao, and Z. Yang, “Visual text generation in the wild,” in ECCV . Springer, 2024, pp. 89–106
2024
Later among the works it cites.
D. Peng, C. Liu, Y. Liu, and L. Jin, “Viteraser: Harnessing the power of vision transformers for scene text removal with segmim pretraining,” in AAAI , 2024, pp. 4468–4477
2024
Later among the works it cites.
OpenAI, “Gpt-4o,” https://openai.com/index/hello-gpt-4o/ , May 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Peng, Z. Yang, J. Zhang, C. Liu, Y. Shi, K. Ding, F. Guo, and L. Jin, “UPOCR: Towards unified pixel-level ocr interface,” in ICML , 2024
2024
Later among the works it cites.
M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M.-H. Yang, and F. S. Khan, “Foundation models defining a new era in vision: A survey and outlook,” TPAMI , 2025
2025
Closest in time.
Y. Du, Z. Chen, Y. Su, C. Jia, and Y.-G. Jiang, “Instruction-guided scene text recognition,” TPAMI , 2025
2025
Closest in time.
X. Yang, Z. Qiao, and Y. Zhou, “IPAD: Iterative, parallel, and diffusion-based network for scene text recognition,” IJCV , 2025
2025
Closest in time.
Y. Zhang, C. Liu, J. Wei, X. Yang, Y. Zhou, C. Ma, and X. Ji, “Linguistics-aware masked image modeling for self-supervised scene text recognition,” in CVPR , 2025
2025
Closest in time.
Z. Yang, D. Peng, Y. Shi, Y. Zhang, C. Liu, and L. Jin, “Predicting the original appearance of damaged historical documents,” in AAAI , vol. 39, no. 9, 2025, pp. 9382–9390
2025
Closest in time.
E. Xie, J. Lyu, D. Wu, H. Shen, and Y. Zhou, “Char-SAM: Turning segment anything model into scene text segmentation annotator with character-level visual prompts,” in ICASSP , 2025
2025
Closest in time.
C. Qu, Y. Zhong, F. Guo, and L. Jin, “Revisiting tampered scene text detection in the era of generative AI,” in AAAI , vol. 39, no. 1, 2025, pp. 694–702
2025
Closest in time.
S. Singh, P. Keserwani, M. Iwamura, and P. P. Roy, “DCDM: Diffusion-conditioned-diffusion model for scene text image super-resolution,” in ECCV . Springer, 2025, pp. 303–320
2025
Closest in time.
C. Liu, Q. Jiang, D. Peng, Y. Kong, J. Zhang, L. Xiong, J. Duan, C. Sun, and L. Jin, “QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a query-aware transformer,” Neurocomputing , vol. 620, p. 129241, 2025
2025
Closest in time.
C. He, Y. Shen, C. Fang, F. Xiao, L. Tang, Y. Zhang, W. Zuo, Z. Guo, and X. Li, “Diffusion models in low-level vision: A survey,” TPAMI , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Qiao, Y. Zhou, J. Wei, W. Wang, Y. Zhang, N. Jiang, H. Wang, and W. Wang, “PIMNet: A parallel, iterative and mimicking network for scene text recognition,” in ACM MM , 2021, pp. 2046–2055
2055
Closest in time.