Fetching the paper…
Reading the bibliography…
Plain text has become a prevalent interface for text-to-image synthesis.
Sahuguet, A., Azavant, F.: Wysiwyg web wrapper factory (w4f) (1999)
1999
Earlier work this paper cites.
Shi, J., Malik, J.: Normalized cuts and image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 22
2000
Earlier work this paper cites.
Reinhard, E., Adhikhmin, M., Gooch, B., Shirley, P.: Color transfer between images. IEEE Computer Graphics and Applications 21
2001
Earlier work this paper cites.
Levin, A., Lischinski, D., Weiss, Y.: Colorization using optimization. SIG (2004)
2004
Earlier work this paper cites.
Tai, Y.-W., Jia, J., Tang, C.-K.: Local color transfer via probabilistic segmentation by expectation-maximization. In: CVPR (2005)
2005
Earlier work this paper cites.
Zhu, X., Goldberg, A.B., Eldawy, M., Dyer, C.R., Strock, B.: A text-to-picture synthesis system for augmenting communication. In: AAAI (2007)
2007
Earlier work this paper cites.
Von Luxburg, U.: A tutorial on spectral clustering. Statistics and computing 17
2007
Earlier work this paper cites.
Witten, I.H., Bainbridge, D., Nichols, D.M.: How to Build a Digital Library, (2009)
2009
Earlier work this paper cites.
Colorado State University, T.A.P.: tutorial: Rich Text Format (RTF) from Microsoft Word - The ACCESS Project (2012). http://accessproject.colostate.edu/udl/modules/word/tut_rtf.cfm
2012
Earlier work this paper cites.
Xu, L., Yan, Q., Jia, J.: A sparse control model for image and video editing. ACM Transactions on Graphics (TOG) 32
2013
Earlier work this paper cites.
Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: CVPR (2015)
2015
Earlier work this paper cites.
Mansimov, E., Parisotto, E., Ba, J.L., Salakhutdinov, R.: Generating images from captions with attention. In: ICLR (2016)
2016
Earlier work this paper cites.
Gatys, L.A., Ecker, A.S., Bethge, M.: Image style transfer using convolutional neural networks. In: CVPR (2016)
2016
Earlier work this paper cites.
Zhang, R., Isola, P., Efros, A.A.: Colorful image colorization. In: ECCV (2016)
2016
Earlier work this paper cites.
Johnson, J., Karpathy, A., Fei-Fei, L.: Densecap: Fully convolutional localization networks for dense captioning. In: CVPR (2016)
2016
Earlier work this paper cites.
Johnson, J., Alahi, A., Fei-Fei, L.: Perceptual losses for real-time style transfer and super-resolution. In: ECCV (2016)
2016
Earlier work this paper cites.
Isola, P., Zhu, J.-Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. In: CVPR (2017)
2017
Earlier work this paper cites.
Zhu, J.-Y., Park, T., Isola, P., Efros, A.A.: Unpaired image-to-image translation using cycle-consistent adversarial networks. In: ICCV (2017)
2017
Earlier work this paper cites.
Luan, F., Paris, S., Shechtman, E., Bala, K.: Deep photo style transfer. In: CVPR (2017)
2017
Earlier work this paper cites.
Zhang, R., Zhu, J.-Y., Isola, P., Geng, X., Lin, A.S., Yu, T., Efros, A.A.: Real-time user-guided image colorization with learned deep priors. ACM Transactions on Graphics (TOG) 9
2017
Earlier work this paper cites.
Johnson, J., Gupta, A., Fei-Fei, L.: Image generation from scene graphs. In: CVPR (2018)
2018
Earlier work this paper cites.
Park, T., Liu, M.-Y., Wang, T.-C., Zhu, J.-Y.: Semantic image synthesis with spatially-adaptive normalization. In: CVPR (2019)
2019
Earlier work this paper cites.
Meng, Y., Wu, W., Wang, F., Li, X., Nie, P., Yin, F., Li, M., Han, Q., Sun, X., Li, J.: Glyce: Glyph-vectors for chinese character representations. NeurIPS 32
2019
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. NeurIPS (2020)
2020
Earlier work this paper cites.
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., Zhou, M.: Layoutlm: Pre-training of text and layout for document image understanding. In: SIGKDD (2020)
2020
Earlier work this paper cites.
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I.: Zero-shot text-to-image generation. In: ICML, pp. 8821–8831 (2021)
2021
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. In: ICLR (2021)
2021
Earlier work this paper cites.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. In: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications (2021)
2021
Earlier work this paper cites.
Choi, J., Kim, S., Jeong, Y., Gwon, Y., Yoon, S.: Ilvr: Conditioning method for denoising diffusion probabilistic models. In: ICCV (2021)
2021
Earlier work this paper cites.
Sun, Z., Li, X., Sun, X., Meng, Y., Ao, X., He, Q., Wu, F., Li, J.: Chinesebert: Chinese pretraining enhanced by glyph and pinyin information. ACL (2021)
2021
Earlier work this paper cites.
Ignat, C.-L., André, L., Oster, G.: Enhancing rich content wikis with real-time collaboration. Concurrency and Computation: Practice and Experience 33
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
2021
Earlier work this paper cites.
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S.K.S., Ayan, B.K., Mahdavi, S.S., Lopes, R.G., Salimans, T., Ho, J., Fleet, D.J., Norouzi, M.: Photorealistic text-to-image diffusion models with deep language understanding. In: NeurIPS (2022)
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., Taigman, Y.: Make-a-scene: Scene-based text-to-image generation with human priors. In: ECCV (2022)
2022
Earlier work this paper cites.
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al
2022
Earlier work this paper cites.
Byeon, M., Park, B., Kim, H., Lee, S., Baek, W., Kim, S.: COYO-700M: Image-Text Pair Dataset. https://github.com/kakaobrain/coyo-dataset (2022)
2022
Earlier work this paper cites.
Ho, J., Saharia, C., Chan, W., Fleet, D.J., Norouzi, M., Salimans, T.: Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research 23
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Nichol, A.Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., Chen, M.: Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In: ICML (2022)
2022
Cited alongside, same era.
Yu, J., Xu, Y., Koh, J.Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B.K., et al.: Scaling autoregressive models for content-rich text-to-image generation. Transactions on Machine Learning Research (2022)
2022
Cited alongside, same era.
Feng, W., He, X., Fu, T.-J., Jampani, V., Akula, A.R., Narayana, P., Basu, S., Wang, X.E., Wang, W.Y.: Training-free structured diffusion guidance for compositional text-to-image synthesis. In: ICLR (2023)
2023
Closest in time.
Couairon, G., Careil, M., Cord, M., Lathuiliere, S., Verbeek, J.: Zero-shot spatial layout conditioning for text-to-image diffusion models. In: ICCV (2023)
2023
Closest in time.
Park, M., Yun, J., Choi, S., Choo, J.: Learning to generate semantic layouts for higher text-image correspondence in text-to-image synthesis. In: ICCV (2023)
2023
Closest in time.
Qu, L., Wu, S., Fei, H., Nie, L., Chua, T.-S.: Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation. In: ACM MM (2023)
2023
Closest in time.
Liu, R., Wu, R., Van Hoorick, B., Tokmakov, P., Zakharov, S., Vondrick, C.: Zero-1-to-3: Zero-shot one image to 3d object. In: ICCV (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Meng, C., Song, Y., Song, J., Wu, J., Zhu, J.-Y., Ermon, S.: Sdedit: Image synthesis and editing with stochastic differential equations. In: ICLR (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., Fleet, D.J.: Video diffusion models. In: NeurIPS (2022)
2022
Cited alongside, same era.
Watson, D., Chan, W., Martin-Brualla, R., Ho, J., Tagliasacchi, A., Norouzi, M.: Novel View Synthesis with Diffusion Models (2022)
2022
Cited alongside, same era.
Li, J., Xu, Y., Lv, T., Cui, L., Zhang, C., Wei, F.: Dit: Self-supervised pre-training for document image transformer. In: MM (2022)
2022
Cited alongside, same era.
Litt, G., Lim, S., Kleppmann, M., Hardenberg, P.: Peritext: A crdt for collaborative rich text editing. Proceedings of the ACM on Human-Computer Interaction (PACMHCI) (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Closest in time.
Tseng, H.-Y., Li, Q., Kim, C., Alsisan, S., Huang, J.-B., Kopf, J.: Consistent view synthesis with pose-guided diffusion models. In: CVPR (2023)
2023
Closest in time.
Patashnik, O., Garibi, D., Azuri, I., Averbuch-Elor, H., Cohen-Or, D.: Localizing object-level shape variations with text-to-image diffusion models. In: ICCV (2023)
2023
Closest in time.
2023
Closest in time.
QI, C., Cun, X., Zhang, Y., Lei, C., Wang, X., Shan, Y., Chen, Q.: Fatezero: Fusing attentions for zero-shot text-based video editing. In: ICCV (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al
2023
Closest in time.
2023
Closest in time.
Bakr, E.M., Sun, P., Shen, X., Khan, F.F., Li, L.E., Elhoseiny, M.: Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models. In: ICCV (2023)
2023
Closest in time.
Hu, Y., Liu, B., Kasai, J., Wang, Y., Ostendorf, M., Krishna, R., Smith, N.A.: Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. In: ICCV (2023)
2023
Closest in time.
Agrawala, M.: Unpredictable Black Boxes are Terrible Interfaces. https://magrawala.substack.com/p/unpredictable-black-boxes-are-terrible (2023)
2023
Closest in time.
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A.H., Chechik, G., Cohen-or, D.: An image is worth one word: Personalizing text-to-image generation using textual inversion. In: ICLR (2023). https://openreview.net/forum?id=NAQvF08TcyG
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Gal, R., Arar, M., Atzmon, Y., Bermano, A.H., Chechik, G., Cohen-Or, D.: Encoder-based domain tuning for fast personalization of text-to-image models. ACM Transactions on Graphics (TOG) 42
2023
Closest in time.
2023
Closest in time.
Tumanyan, N., Geyer, M., Bagon, S., Dekel, T.: Plug-and-play diffusion features for text-driven image-to-image translation. In: CVPR (2023)
2023
Closest in time.
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., Goldstein, T.: Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. In: NeurIPS (2023)
2023
Closest in time.
OpenAI: GPT-4 Technical Report (2023)
2023
Closest in time.
2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.-Y., Dollar, P., Girshick, R.: Segment anything. In: ICCV (2023)
2023
Closest in time.
2023
Closest in time.
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., Zeng, Z., Zhang, H., Li, F., Yang, J., Li, H., Jiang, Q., Zhang, L.: Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks (2024)
2023
Closest in time.
Phung, Q., Ge, S., Huang, J.-B.: Grounded text-to-image synthesis with attention refocusing. In: CVPR (2024)
2024
Closest in time.
Feng, W., Zhu, W., Fu, T.-j., Jampani, V., Akula, A., He, X., Basu, S., Wang, X.E., Wang, W.Y.: Layoutgpt: Compositional visual planning and generation with large language models. (2024)
2024
Closest in time.
2024
Closest in time.
Huang, K., Sun, K., Xie, E., Li, Z., Liu, X.: T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation. In: NeurIPS (2024)
2024
Closest in time.
Patel, M., Gokhale, T., Baral, C., Yang, Y.: Conceptbed: Evaluating concept learning abilities of text-to-image diffusion models. In: AAAI (2024)
2024
Closest in time.
2024
Closest in time.
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: SDXL: Improving latent diffusion models for high-resolution image synthesis. In: ICLR (2024)
2024
Closest in time.
Ju, X., Zeng, A., Bian, Y., Liu, S., Xu, Q.: Pnp inversion: Boosting diffusion-based editing with 3 lines of code. In: ICLR (2024)
2024
Closest in time.