Fetching the paper…
Reading the bibliography…
Modern text-to-vision generative models often hallucinate when the prompt describing the scene to be generated is underspecified.
Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The caltech-ucsd birds-200-2011 datase. Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)
2011
Earlier work this paper cites.
2014
Earlier work this paper cites.
Secretariat, G.: Global biodiversity information facility backbone taxonomy. https://doi.org/10.15468/39omei (2014)
2014
Earlier work this paper cites.
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X.: Improved techniques for training gans. In: NeurIPS (2016)
2016
Earlier work this paper cites.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Johnson, J., Douze, M., Jégou, H.: Billion-scale similarity search with GPUs. In: IEEE Transactions on Big Data (2019)
2019
Earlier work this paper cites.
Karras, T., Laine, S., Aila, T.: A Style-Based Generator Architecture for Generative Adversarial Networks. In: CVPR (2019)
2019
Earlier work this paper cites.
Guu, K., Lee, K., Tung, Z., Pasupat, P., Chang, M.W.: Realm: Retrieval-augmented language model pre-training. In: ICML (2020)
2020
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. In: NeurIPS (2020)
2020
Earlier work this paper cites.
Karras, T., Aittala, M., Hellsten, J., Laine, S., Lehtinen, J., Aila, T.: Training Generative Adversarial Networks with Limited Data. In: NeurIPS (2020)
2020
Earlier work this paper cites.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., tau Yih, W., Rocktäschel, T., Riedel, S., Kiela, D.: Retrieval-augmented generation for knowledge-intensive nlp tasks. In: NeurIPS (2020)
2020
Earlier work this paper cites.
Dhariwal, P., Nichol, A.Q.: Diffusion models beat gans on image synthesis. In: NeurIPS (2021)
2021
Earlier work this paper cites.
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J.: Measuring massive multitask language understanding. In: ICLR (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., Lischinski, D.: StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. In: ICCV (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML (2021)
2021
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR (2021)
2021
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. In: ICLR (2021)
2021
Earlier work this paper cites.
Wu, Z., Lischinski, D., Shechtman, E.: StyleSpace Analysis: Disentangled Controls for StyleGAN Image Generation. In: CVPR (2021)
2021
Earlier work this paper cites.
Blattmann, A., Rombach, R., Oktay, K., Müller, J., Ommer, B.: Semi-parametric neural image synthesis. In: NeurIPS (2022)
2022
Cited alongside, same era.
Chattopadhyay, A., Slocum, S., Haeffele, B.D., Vidal, R., Geman, D.: Interpretable by design: Learning predictors by composing interpretable queries. In: IEEE TPAMI (2022)
2022
Cited alongside, same era.
Competition, G.U.I.E.: Guie laion-5b dataset. Kaggle (2022)
2022
Cited alongside, same era.
Hu, E.J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: ICLR (2022)
2022
Cited alongside, same era.
Kim, G., Kwon, T., Ye, J.C.: Diffusionclip: Text-guided diffusion models for robust image manipulation. In: CVPR (2022)
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
Luo, J., Wang, Z., Wu, C.H., Huang, D., Torre, F.D.L.: Zero-shot model diagnosis. In: CVPR (2023)
2023
Closest in time.
Luo, Z., Chen, D., Zhang, Y., Huang, Y., Wang, L., Shen, Y., Zhao, D., Zhou, J., Tan, T.: Videofusion: Decomposed diffusion models for high-quality video generation. In: CVPR (2023)
2023
Closest in time.
Nguyen, T., Gadre, S.Y., Ilharco, G., Oh, S., Schmidt, L.: Improving multimodal datasets with image captioning. In: NeurIPS Datasets and Benchmarks Track (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., Jitsev, J.: Laion-5b: An open large-scale dataset for training next generation image-text models. In: NeurIPS Datasets and Benchmarks Track (2022)
2022
Cited alongside, same era.
Yu, J., Xu, Y., Koh, J.Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B.K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., Wu, Y.: Scaling autoregressive models for content-rich text-to-image generation. In: TMLR (2022)
2022
Cited alongside, same era.
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., Manassra, W., Dhariwal, P., Chu, C., Jiao, Y.: Improving image generation with better captions. https://cdn.openai.com/papers/dall-e-3.pdf (2023)
2023
Cited alongside, same era.
Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., Caliskan, A.: Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. In: FAccT (2023)
2023
Cited alongside, same era.
Chattopadhyay, A., Chan, K.H.R., Haeffele, B.D., Geman, D., Vidal, R.: Variational information pursuit for interpretable predictions. In: ICLR (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Cho, J., Zala, A., Bansal, M.: Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In: ICCV (2023)
2023
Cited alongside, same era.
2023
Closest in time.
OpenAI: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
2023
Closest in time.
2023
Closest in time.
Poole, B., Jain, A., Barron, J.T., Mildenhall, B.: Dreamfusion: Text-to-3d using 2d diffusion. In: ICLR (2023)
2023
Closest in time.
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., tau Yih, W.: Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv: 2301.12652 (2023)
2023
Closest in time.
Shonenkov, A., Konstantinov, M., Bakshandaeva, D., Schuhmann, C., Ivanova, K., Klokova, N.: Deepfloyd-if. Hugging Face (2023)
2023
Closest in time.
2023
Closest in time.
Sterling, S.: Zeroscope v2. Hugging Face (2023)
2023
Closest in time.
2023
Closest in time.
Wang, Z., Lu, C., Wang, Y., Bao, F., Li, C., Su, H., Zhu, J.: Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. In: NeurIPS (2023)
2023
Closest in time.
Yasunaga, M., Aghajanyan, A., Shi, W., James, R., Leskovec, J., Liang, P., Lewis, M., Zettlemoyer, L., tau Yih, W.: Retrieval-augmented multimodal language modeling. In: ICML (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Wang, Y.O., Chung, Y., Wu, C.H., De la Torre, F.: Domain gap embeddings for generative dataset augmentation. In: CVPR (2024)
2024
Closest in time.