Fetching the paper…
Reading the bibliography…
Recently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images.
Unrealtext: Synthesizing realistic scene text images from the unreal world
Long, S.; and Yao, C. 2020 · 2003
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Gupta, A.; Vedaldi, A.; and Zisserman, A. 2016 · 2016
Earlier work this paper cites.
Icdar2017 competition on reading chinese text in the wild (rctw-17)
Shi, B.; Yao, C.; Liao, M.; Yang, M.; Xu, P.; Cui, L.; Belongie, S.; Lu, S.; and Bai, X. 2017 · 2017
Earlier work this paper cites.
Editing text in the wild
Wu, L.; Zhang, C.; Liu, J.; Han, J.; Liu, J.; Ding, E.; and Bai, X. 2019 · 2019
Earlier work this paper cites.
Spatial fusion gan for image synthesis
Zhan, F.; Zhu, H.; and Lu, S. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
SynthText3D: synthesizing scene text images from 3D virtual worlds
Liao, M.; Song, B.; Long, S.; He, M.; Yao, C.; and Bai, X. 2020 · 2020
Earlier work this paper cites.
STEFANN: scene text editor using font adaptive neural network
Roy, P.; Bhattacharya, S.; Ghosh, S.; and Pal, U. 2020 · 2020
Earlier work this paper cites.
Swaptext: Image based texts transfer in scenes
Yang, Q.; Huang, J.; and Lin, W. 2020 · 2020
Earlier work this paper cites.
Bau, D.; Andonian, A.; Cui, A.; Park, Y.; Jahanian, A.; Oliva, A.; and Torralba, A. 2021 · 2021
Earlier work this paper cites.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Unifying multimodal transformer for bi-directional image and text generation
Huang, Y.; Xue, H.; Liu, B.; and Lu, Y. 2021 · 2021
Earlier work this paper cites.
Rewritenet: Realistic scene text image generation via editing text in real-world image
Lee, J.; Kim, Y.; Kim, S.; Yim, M.; Shin, S.; Lee, G.; and Park, S. 2021 · 2021
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.-Y.; and Ermon, S. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
DG-Font: Deformable Generative Networks for Unsupervised Font Generation
Xie, Y.; Chen, X.; Sun, L.; and Lu, Y. 2021 · 2021
Cited alongside, same era.
Scene Text Transfer for Cross-Language
Zhang, L.; Chen, X.; Xie, Y.; and Lu, Y. 2021 · 2021
Cited alongside, same era.
Clipscore
2022 · 2022
Cited alongside, same era.
Blended diffusion for text-driven editing of natural images
Avrahami, O.; Lischinski, D.; and Fried, O. 2022 · 2022
Cited alongside, same era.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; Catanzaro, B.; et al. 2022 · 2022
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
DeepFloyd IF
2023 · 2023
Closest in time.
Align your latents: High-resolution video synthesis with latent diffusion models
Blattmann, A.; Rombach, R.; Ling, H.; Dockhorn, T.; Kim, S. W.; Fidler, S.; and Kreis, K. 2023 · 2023
Closest in time.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T.; Holynski, A.; and Efros, A. A. 2023 · 2023
Closest in time.
Preserve your own correlation: A noise prior for video diffusion models
Ge, S.; Nah, S.; Liu, G.; Poon, T.; Tao, A.; Catanzaro, B.; Jacobs, D.; Huang, J.-B.; Liu, M.-Y.; and Balaji, Y. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diff-Font: Diffusion Model for Robust One-Shot Font Generation
He, H.; Chen, X.; Wang, C.; Liu, J.; Du, B.; Tao, D.; and Qiao, Y. 2022 · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Hertz, A.; Mokady, R.; Tenenbaum, J.; Aberman, K.; Pritch, Y.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; et al. 2022 · 2022
Cited alongside, same era.
Repaint: Inpainting using denoising diffusion probabilistic models
Lugmayr, A.; Danelljan, M.; Romero, A.; Yu, F.; Timofte, R.; and Van Gool, L. 2022 · 2022
Cited alongside, same era.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Nichol, A. Q.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; Mcgrew, B.; Sutskever, I.; and Chen, M. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Kawar, B.; Zada, S.; Lang, O.; Tov, O.; Chang, H.; Dekel, T.; Mosseri, I.; and Irani, M. 2023 · 2023
Closest in time.
Textstylebrush: Transfer of text aesthetics from a single example
Krishnan, P.; Kovvuri, R.; Pang, G.; Vassilev, B.; and Hassner, T. 2023 · 2023
Closest in time.
More control for free! image synthesis with semantic diffusion guidance
Liu, X.; Park, D. H.; Azadi, S.; Zhang, G.; Chopikyan, A.; Hu, Y.; Shi, H.; Rohrbach, A.; and Darrell, T. 2023 · 2023
Closest in time.
GlyphDraw: Learning to Draw Chinese Characters in Image Synthesis Models Coherently
Ma, J.; Zhao, M.; Chen, C.; Wang, R.; Niu, D.; Lu, H.; and Lin, X. 2023 · 2023
Closest in time.
Null-text inversion for editing real images using guided diffusion models
Mokady, R.; Hertz, A.; Aberman, K.; Pritch, Y.; and Cohen-Or, D. 2023 · 2023
Closest in time.
Mou, C.; Wang, X.; Xie, L.; Zhang, J.; Qi, Z.; Shan, Y.; and Qie, X. 2023 · 2023
Closest in time.
Make-A-Video: Text-to-Video Generation without Text-Video Data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; Parikh, D.; Gupta, S.; and Taigman, Y. 2023 · 2023
Closest in time.
Weakly supervised scene text generation for low-resource languages
Xie, Y.; Chen, X.; Zhan, H.; Shivakumara, P.; Yin, B.; Liu, C.; and Lu, Y. 2023 · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Closest in time.