Fetching the paper…
Reading the bibliography…
To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model's ability to capture the true semantic representation of visual scenes.
Stacked quantizers for compositional vector compression
Martinez, J., Hoos, H. H., and Little, J. J · 2014
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Figr: Few-shot image generation with reptile
Clouâtre, L. and Demers, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Deepsvg: A hierarchical generative network for vector graphics animation
Carlier, A., Danelljan, M., Alahi, A., and Timofte, R · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Earlier work this paper cites.
Differentiable vector graphics rasterization for editing and learning
Li, T.-M., Lukáč, M., Gharbi, M., and Ragan-Kelley, J · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Dall-e: Creating images from text
Reddy, M. D. M., Basha, M. S. M., Hari, M. M. C., and Penchalaiah, M. N · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Draft-and-revise: Effective image generation with contextual rq-transformer
Lee, D., Kim, C., Kim, S., Cho, M., and HAN, W. S · 2022
Cited alongside, same era.
Towards layer-wise image vectorization
Ma, X., Zhou, Y., Xu, X., Sun, B., Filev, V., Orlov, N., Fu, Y., and Shi, H · 2022
Cited alongside, same era.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
Ghosal, D., Majumder, N., Mehrish, A., and Poria, S · 2023
Later among the works it cites.
Robust semantic communications with masked vq-vae enabled codebook
Hu, Q., Zhang, G., Qin, Z., Cai, Y., Yu, G., and Li, G. Y · 2023
Later among the works it cites.
Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models
Jain, A., Xie, A., and Abbeel, P · 2023
Later among the works it cites.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
Later among the works it cites.
Towards expert-level medical question answering with large language models
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Hou, L., Clark, K., Pfohl, S., Cole-Lewis, H., Neal, D., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Aesthetic text logo synthesis via content-aware layout inferring
Wang, Y., Pu, G., Luo, W., Wang, Y., Xiong, P., Kang, H., and Lian, Z · 2022
Cited alongside, same era.
Nüwa: Visual synthesis pre-training for neural visual world creation
Wu, C., Liang, J., Ji, L., Yang, F., Fang, Y., Jiang, D., and Duan, N · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Cited alongside, same era.
Efficient-vqgan: Towards high-resolution image generation with efficient vision transformers
Cao, S., Yin, Y., Huang, L., Liu, Y., Zhao, X., Zhao, D., and Huang, K · 2023
Cited alongside, same era.
Later among the works it cites.
Marvel: Raster gray-level manga vectorization via primitive-wise deep reinforcement learning
Su, H., Liu, X., Niu, J., Cui, J., Wan, J., Wu, X., and Wang, N · 2023
Later among the works it cites.
Generative pretraining in multimodality
Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Svgdreamer: Text guided svg generation with diffusion model
Xing, X., Zhou, H., Wang, C., Zhang, J., Xu, D., and Yu, Q · 2023
Later among the works it cites.
Language model beats diffusion–tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al · 2023
Later among the works it cites.
Beyond pixels: Exploring human-readable svg generation for simple images with vision language models
Zhang, T., Liu, H., Zhang, P., Cheng, Y., and Wang, H · 2023
Later among the works it cites.
Human preference score: Better aligning text-to-image models with human preference
Wu, X., Sun, K., Zhu, F., Zhao, R., and Li, H · 2096
Closest in time.