Fetching the paper…
Reading the bibliography…
Text-to-image generation (TTI) refers to the usage of models that could process text input and generate high fidelity images based on text descriptions.
K. Pearson, “Liii. on lines and planes of closest fit to systems of points in space,” The London, Edinburgh, and Dublin philosophical magazine and journal of science , vol. 2, no. 11, pp. 559–572, 1901
1901
Earlier work this paper cites.
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel, “Plug-and-play diffusion features for text-driven image-to-image translation,” in CVPR , 2023, pp. 1921–1930
1930
Earlier work this paper cites.
J. Cho, J. Lei, H. Tan, and M. Bansal, “Unifying vision-and-language tasks via text generation,” in ICML . PMLR, 2021, pp. 1931–1942
1942
Earlier work this paper cites.
“Comptes rendus hebdomadaires des seances de l’academie des sciences,” 1835–1965
1965
Earlier work this paper cites.
B. D. Anderson, “Reverse-time diffusion equation models,” Stochastic Processes and their Applications , vol. 12, no. 3, pp. 313–326, 1982
1982
Earlier work this paper cites.
D. Dowson and B. Landau, “The fréchet distance between multivariate normal distributions,” Journal of multivariate analysis , vol. 12, no. 3, pp. 450–455, 1982
1982
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
C. B. Do, “The multivariate gaussian distribution,” Section Notes, Lecture on Machine Learning, CS , vol. 229, 2008
2008
Earlier work this paper cites.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009, pp. 248–255
2009
Earlier work this paper cites.
P. Vincent, “A connection between score matching and denoising autoencoders,” Neural Computation , vol. 23, no. 7, pp. 1661–1674, 2011
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic backpropagation and approximate inference in deep generative models,” in ICML . PMLR, 2014, pp. 1278–1286
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” NeurIPS , vol. 27, 2014
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder–decoder approaches,” in Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation . Association for Computational Linguistics, 2014, p. 103
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
J. Gauthier, “Conditional generative adversarial nets for convolutional face generation,” Class project for Stanford CS231N: convolutional neural networks for visual recognition, Winter semester , vol. 2014, no. 5, p. 2, 2014
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
“Deep learning,” Nature Cell Biology , vol. 521, no. 7553, pp. 436–444, May 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Dong, C. C. Loy, K. He, and X. Tang, “Image super-resolution using deep convolutional networks,” TPAMI , vol. 38, no. 2, pp. 295–307, 2015
2015
Earlier work this paper cites.
E. L. Denton, S. Chintala, R. Fergus et al. , “Deep generative image models using a laplacian pyramid of adversarial networks,” NeurIPS , vol. 28, 2015
2015
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in ICML . PMLR, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, “Generative adversarial text to image synthesis,” in ICML . PMLR, 2016, pp. 1060–1069
2016
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
M.-Y. Liu and O. Tuzel, “Coupled generative adversarial networks,” NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, “Generative adversarial text to image synthesis,” in ICML . PMLR, 2016, pp. 1060–1069
2016
Earlier work this paper cites.
S. E. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee, “Learning what and where to draw,” NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in 54th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics (ACL), 2016, pp. 1715–1725
2016
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in ICLR , 2016
2016
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in ICCV , 2017, pp. 2223–2232
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
J. Wu, “Introduction to convolutional neural networks,” National Key Lab for Novel Software Technology. Nanjing University. China , vol. 5, no. 23, p. 495, 2017
2017
Earlier work this paper cites.
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang et al. , “Photo-realistic single image super-resolution using a generative adversarial network,” in CVPR , 2017, pp. 4681–4690
2017
Earlier work this paper cites.
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas, “Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks,” in ICCV , 2017, pp. 5907–5915
2017
Earlier work this paper cites.
A. Odena, C. Olah, and J. Shlens, “Conditional image synthesis with auxiliary classifier gans,” in ICML . PMLR, 2017, pp. 2642–2651
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. T. Rolfe, “Discrete variational autoencoders,” in ICLR , 2017
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He, “Attngan: Fine-grained text to image generation with attentional generative adversarial networks,” in CVPR , 2018, pp. 1316–1324
2018
Earlier work this paper cites.
A. Brock, J. Donahue, and K. Simonyan, “Large scale gan training for high fidelity natural image synthesis,” in ICLR , 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
Z. Zhang, Y. Xie, and L. Yang, “Photographic text-to-image synthesis with a hierarchically-nested adversarial network,” in CVPR , 2018, pp. 6199–6208
2018
Earlier work this paper cites.
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas, “Stackgan++: Realistic image synthesis with stacked generative adversarial networks,” TPAMI , vol. 41, no. 8, pp. 1947–1962, 2018
2018
Earlier work this paper cites.
J. Johnson, A. Gupta, and L. Fei-Fei, “Image generation from scene graphs,” in CVPR , 2018, pp. 1219–1228
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu, “Semantic image synthesis with spatially-adaptive normalization,” in CVPR , 2019, pp. 2337–2346
2019
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
M. Zhu, P. Pan, W. Chen, and Y. Yang, “Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis,” in CVPR , 2019, pp. 5802–5810
2019
Earlier work this paper cites.
T. Qiao, J. Zhang, D. Xu, and D. Tao, “Mirrorgan: Learning text-to-image generation by redescription,” in CVPR , 2019, pp. 1505–1514
2019
Earlier work this paper cites.
M. Cha, Y. L. Gwon, and H. Kung, “Adversarial learning of semantic relevance in text to image synthesis,” in AAAI , vol. 33, no. 01, 2019, pp. 3272–3279
2019
Earlier work this paper cites.
Y. Yang, H. Li, X. Li, Q. Zhao, J. Wu, and Z. Lin, “Sognet: Scene overlap graph network for panoptic segmentation,” in AAAI , 2019
2019
Earlier work this paper cites.
L. Liu, M. Muelly, J. Deng, T. Pfister, and L.-J. Li, “Generative modeling for small-data object detection,” in ICCV , 2019, pp. 6073–6081
2019
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. Dong, N. Yang, W. Wang, F. Wei, X. Liu, Y. Wang, J. Gao, M. Zhou, and H.-W. Hon, “Unified language model pre-training for natural language understanding and generation,” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “Electra: Pre-training text encoders as discriminators rather than generators,” in ICLR , 2019
2019
Earlier work this paper cites.
Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, “Ernie: Enhanced language representation with informative entities,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 1441–1451
2019
Earlier work this paper cites.
A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, “Panoptic segmentation,” in CVPR , 2019, pp. 9404–9413
2019
Earlier work this paper cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in CVPR , 2019, pp. 10 502–10 511
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. He, Z. Zhang, H. Zhang, Z. Zhang, J. Xie, and M. Li, “Bag of tricks for image classification with convolutional neural networks,” in CVPR , 2019, pp. 558–567
2019
Earlier work this paper cites.
L. Gao, D. Chen, J. Song, X. Xu, D. Zhang, and H. T. Shen, “Perceptual pyramid adversarial networks for text-to-image synthesis,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 8312–8319
2019
Earlier work this paper cites.
G. Yin, B. Liu, L. Sheng, N. Yu, X. Wang, and J. Shao, “Semantics disentangling for text-to-image generation,” in CVPR , 2019, pp. 2327–2336
2019
Earlier work this paper cites.
H. Liu, X. Shen, F. Shang, F. Ge, and F. Wang, “Cu-net: Cascaded u-net with loss weighted sampling for brain tumor segmentation,” in Multimodal Brain Image Analysis and Mathematical Foundations of Computational Anatomy: 4th International Workshop, MBIA 2019, and 7th International Workshop, MFCA 2019, Held in Conjunction with MICCAI 2019, Shenzhen, China, October 17, 2019, Proceedings 4 . Springer, 2019, pp. 102–111
2019
Earlier work this paper cites.
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis, “Generalization through memorization: Nearest neighbor language models,” in ICLR , 2019
2019
Earlier work this paper cites.
W. Li, P. Zhang, L. Zhang, Q. Huang, X. He, S. Lyu, and J. Gao, “Object-driven text-to-image synthesis via adversarial training,” in CVPR , 2019, pp. 12 174–12 182
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” JMLR , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in CVPR , 2020, pp. 8110–8119
2020
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Song, S. Garg, J. Shi, and S. Ermon, “Sliced score matching: A scalable approach to density and score estimation,” in Uncertainty in Artificial Intelligence . PMLR, 2020, pp. 574–584
2020
Earlier work this paper cites.
T. Pang, K. Xu, C. Li, Y. Song, S. Ermon, and J. Zhu, “Efficient learning of generative models via finite-difference score matching,” NeurIPS , vol. 33, pp. 19 175–19 188, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language representations,” in ICLR , 2020
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV . Springer, 2020, pp. 213–229
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang et al. , “Gpt-neox-20b: An open-source autoregressive language model,” in Proceedings of BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models , 2022, pp. 95–136
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 8110–8119
2020
Cited alongside, same era.
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang, “Retrieval augmented language model pre-training,” in ICML . PMLR, 2020, pp. 3929–3938
2020
Cited alongside, same era.
E. Xie, P. Sun, X. Song, W. Wang, X. Liu, D. Liang, C. Shen, and P. Luo, “Polarmask: Single shot instance segmentation with polar representation,” in CVPR , 2020, pp. 12 193–12 202
2020
Cited alongside, same era.
M. J. Chong and D. Forsyth, “Effectively unbiased fid and inception score and where to find them,” in CVPR , 2020, pp. 6070–6079
2020
Cited alongside, same era.
K. Schwarz, Y. Liao, M. Niemeyer, and A. Geiger, “Graf: Generative radiance fields for 3d-aware image synthesis,” NeurIPS , vol. 33, pp. 20 154–20 166, 2020
2020
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in ICML . PMLR, 2021, pp. 8821–8831
2021
Cited alongside, same era.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” NeurIPS , vol. 34, pp. 8780–8794, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” in ICML . PMLR, 2022, pp. 18 332–18 346
2022
Later among the works it cites.
S. Welleck, J. Liu, X. Lu, H. Hajishirzi, and Y. Choi, “Naturalprover: Grounded mathematical proof generation with language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 4913–4927, 2022
2022
Later among the works it cites.
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,” NeurIPS , vol. 35, pp. 32 897–32 912, 2022
2022
Later among the works it cites.
H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,” in ICLR , 2022
2022
Later among the works it cites.
Y. Li, F. Liang, L. Zhao, Y. Cui, W. Ouyang, J. Shao, F. Yu, and J. Yan, “Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm,” in ICLR , 2022
2022
Later among the works it cites.
J. Chen, H. Guo, K. Yi, B. Li, and M. Elhoseiny, “Visualgpt: Data-efficient adaptation of pretrained language models for image captioning,” in CVPR , 2022, pp. 18 030–18 040
2022
Later among the works it cites.
T. Wang, W. Jiang, Z. Lu, F. Zheng, R. Cheng, C. Yin, and P. Luo, “Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix,” in ICML . PMLR, 2022, pp. 22 680–22 690
2022
Later among the works it cites.
Z. Li, Z. Geng, Z. Kang, W. Chen, and Y. Yang, “Eliminating gradient conflict in reference-based line-art colorization,” in ECCV , 2022
2022
Later among the works it cites.
R. Gal, O. Patashnik, H. Maron, A. H. Bermano, G. Chechik, and D. Cohen-Or, “Stylegan-nada: Clip-guided domain adaptation of image generators,” ACM Transactions on Graphics , vol. 41, no. 4, pp. 1–13, 2022
2022
Later among the works it cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS , vol. 35, pp. 25 278–25 294, 2022
2022
Later among the works it cites.
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset , 2022
2022
Later among the works it cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022
Later among the works it cites.
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, “Cascaded diffusion models for high fidelity image generation,” JMLR , vol. 23, no. 1, pp. 2249–2281, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Blattmann, R. Rombach, K. Oktay, J. Müller, and B. Ommer, “Semi-parametric neural image synthesis,” NeurIPS , vol. 11, 2022
2022
Later among the works it cites.
S. Sheynin, O. Ashual, A. Polyak, U. Singer, O. Gafni, E. Nachmani, and Y. Taigman, “knn-diffusion: Image generation via large-scale retrieval,” in The Eleventh ICLR , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark et al. , “Improving language models by retrieving from trillions of tokens,” in ICML . PMLR, 2022, pp. 2206–2240
2022
Later among the works it cites.
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark et al. , “Improving language models by retrieving from trillions of tokens,” in ICML . PMLR, 2022, pp. 2206–2240
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,” Advances in Neural Information Processing Systems , vol. 35, pp. 5775–5787, 2022
2022
Later among the works it cites.
Z. Chen, X. Tan, K. Wang, S. Pan, D. Mandic, L. He, and S. Zhao, “Infergrad: Improving diffusion models for vocoder by considering inference in training,” in ICASSP . IEEE, 2022, pp. 8432–8436
2022
Later among the works it cites.
W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood, “Flexible diffusion modeling of long videos,” NeurIPS , vol. 35, pp. 27 953–27 965, 2022
2022
Later among the works it cites.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni et al. , “Make-a-video: Text-to-video generation without text-video data,” in The Eleventh ICLR , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Wu and C. Zheng, “Learning to generate 3d shapes from a single example,” ACM Transactions on Graphics , vol. 41, no. 6, pp. 1–19, 2022
2022
Later among the works it cites.
X. Zeng, A. Vahdat, F. Williams, Z. Gojcic, O. Litany, S. Fidler, and K. Kreis, “Lion: Latent point diffusion models for 3d shape generation,” in NeurIPS , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
K.-H. Hui, R. Li, J. Hu, and C.-W. Fu, “Neural wavelet-domain diffusion for 3d shape generation,” in SIGGRAPH Asia 2022 Conference Papers , 2022, pp. 1–9
2022
Later among the works it cites.
R. Wang, Y. Yang, and D. Tao, “Art-point: Improving rotation robustness of point cloud classifiers via adversarial rotation,” in CVPR , 2022
2022
Later among the works it cites.
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” 2023
2023
Closest in time.
Y. Zhang, N. Huang, F. Tang, H. Huang, C. Ma, W. Dong, and C. Xu, “Inversion-based style transfer with diffusion models,” in CVPR , 2023, pp. 10 146–10 156
2023
Closest in time.
B. Yang, S. Gu, B. Zhang, T. Zhang, X. Chen, X. Sun, D. Chen, and F. Wen, “Paint by example: Exemplar-based image editing with diffusion models,” in CVPR , 2023, pp. 18 381–18 391
2023
Closest in time.
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in CVPR , 2023, pp. 22 500–22 510
2023
Closest in time.
2023
Closest in time.
M. Tao, B.-K. Bao, H. Tang, and C. Xu, “Galip: Generative adversarial clips for text-to-image synthesis,” in CVPR , 2023, pp. 14 214–14 223
2023
Closest in time.
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park, “Scaling up gans for text-to-image synthesis,” in CVPR , 2023, pp. 10 124–10 134
2023
Closest in time.
M. Yeshasvi, P. Kayal, and T. Subetha, “A survey on text description to image generation using gan,” in Soft Computing for Security Applications . Springer Nature Singapore, 2023, pp. 665–675
2023
Closest in time.
F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” TPAMI , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Zhou, B. Liu, Y. Zhu, X. Yang, C. Chen, and J. Xu, “Shifted diffusion for text-to-image generation,” in CVPR , 2023, pp. 10 157–10 166
2023
Closest in time.
2023
Closest in time.
O. Avrahami, T. Hayes, O. Gafni, S. Gupta, Y. Taigman, D. Parikh, D. Lischinski, O. Fried, and X. Yin, “Spatext: Spatio-textual representation for controllable image generation,” in CVPR , 2023, pp. 18 370–18 380
2023
Closest in time.
Z. Feng, Z. Zhang, X. Yu, Y. Fang, L. Li, X. Chen, Y. Lu, J. Liu, W. Yin, S. Feng et al. , “Ernie-vilg 2.0: Improving text-to-image diffusion model with knowledge-enhanced mixture-of-denoising-experts,” in CVPR , 2023, pp. 10 135–10 145
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
L. Ruan, Y. Ma, H. Yang, H. He, B. Liu, J. Fu, N. J. Yuan, Q. Jin, and B. Guo, “Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,” in CVPR , 2023, pp. 10 219–10 228
2023
Closest in time.
2023
Closest in time.
S. Yu, K. Sohn, S. Kim, and J. Shin, “Video probabilistic diffusion models in projected latent space,” in CVPR , 2023, pp. 18 456–18 466
2023
Closest in time.
H. Ni, C. Shi, K. Li, S. X. Huang, and M. R. Min, “Conditional image-to-video generation with latent flow diffusion models,” in CVPR , 2023, pp. 18 444–18 455
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Sun, X. Wang, L. Wang, X. Li, Y. Zhang, H. Zhang, and Y. Liu, “Next3d: Generative neural texture rasterization for 3d-aware head avatars,” in CVPR , 2023, pp. 20 991–21 002
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
T. Wang, B. Zhang, T. Zhang, S. Gu, J. Bao, T. Baltrusaitis, J. Shen, D. Chen, F. Wen, Q. Chen et al. , “Rodin: A generative model for sculpting 3d digital avatars using diffusion,” in CVPR , 2023, pp. 4563–4573
2023
Closest in time.
J. R. Shue, E. R. Chan, R. Po, Z. Ankner, J. Wu, and G. Wetzstein, “3d neural field generation using triplane diffusion,” in CVPR , 2023, pp. 20 875–20 886
2023
Closest in time.
Z. Liu, Y. Feng, M. J. Black, D. Nowrouzezahrai, L. Paull, and W. Liu, “Meshdiffusion: Score-based generative 3d mesh modeling,” in ICLR , 2023
2023
Closest in time.
2023
Closest in time.