Fetching the paper…
Reading the bibliography…
In this work, we explore a cost-effective framework for multilingual image generation.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 1910
Earlier work this paper cites.
Ccmatrix: Mining billions of high-quality parallel sentences on the web, 2020
Schwenk, H., Wenzek, G., Edunov, S., Grave, E., and Joulin, A · 1911
Earlier work this paper cites.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer, 2021
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Dollár, P · 2015
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision, 2021
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Cross-lingual and multilingual clip
Carlsson, F., Eisen, P., Rekathati, F., and Sahlgren, M · 2022
Earlier work this paper cites.
Altclip: Altering the language encoder in clip for extended language capabilities
Chen, Z., Liu, G., Zhang, B.-W., Ye, F., Yang, Q., and Wu, L · 2022
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
NLLB Team, Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Mejia-Gonzalez, G., Hansanti, P., Hoffman, J., Jarrett, S., Sadagopan, K. R., Rowe, D., Spruit, S., Tran, C., Andrews, P., Ayan, N. F., Bhosale, S., Edunov, S., Fan, A., Gao, C., Goswami, V., Guzmán, F., Koehn, P., Mourachko, A., Ropers, C., Saleem, S., Schwenk, H., and Wang, J · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Journeydb: A benchmark for generative image understanding, 2023
Pan, J., Sun, K., Ge, Y., Li, H., Duan, H., Wu, X., Zhang, R., Zhou, A., Qin, Z., Wang, Y., Dai, J., Qiao, Y., and Li, H · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Gluegen: Plug and play multi-modal encoders for x-to-image generation
Qin, C., Yu, N., Xing, C., Zhang, S., Chen, Z., Ermon, S., Fu, Y., Xiong, C., and Xu, R · 2023
Later among the works it cites.
Mvdream: Multi-view diffusion for 3d generation
Shi, Y., Wang, P., Ye, J., Mai, L., Li, K., and Yang, X · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Crossmodal-3600: A massively multilingual multimodal evaluation dataset
Thapliyal, A. V., Pont-Tuset, J., Chen, X., and Soricut, R · 2022
Cited alongside, same era.
Fengshenbang 1.0: Being the foundation of chinese cognitive intelligence
Zhang, J., Gan, R., Wang, J., Zhang, Y., Zhang, L., Yang, P., Gao, X., Wu, Z., Dong, X., He, J., Zhuo, J., Yang, Q., Huang, Y., Li, X., Wu, Y., Lu, J., Zhu, X., Chen, W., Han, T., Pan, K., Wang, R., Wang, H., Wu, X., Zeng, Z., and Chen, C · 2022
Cited alongside, same era.
Kandinsky 3.0 technical report
Arkhipkin, V., Filatov, A., Vasilev, V., Maltseva, A., Azizov, S., Pavlov, I., Agafonova, J., Kuznetsov, A., and Dimitrov, D · 2023
Cited alongside, same era.
Efficient diffusion training via min-snr weighting strategy
Hang, T., Gu, S., Li, C., Bao, J., Chen, D., Hu, H., Geng, X., and Guo, B · 2023
Cited alongside, same era.
Lu, G., Guo, Y., Han, J., Niu, M., Zeng, Y., Xu, S., Huang, Z., Zhong, Z., Zhang, W., and Xu, H · 2023
Cited alongside, same era.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., et al · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Closest in time.
Clip-vit-h-14-frozen-xlm-roberta-large-laion5b-s13b-b90k, 2023
LAION · 2024
Closest in time.
Evaluating text-to-visual generation with image-to-text generation
Lin, Z., Pathak, D., Li, B., Li, J., Xia, X., Neubig, G., Zhang, P., and Ramanan, D · 2024
Closest in time.
Recent advances in implicit representation-based 3d shape generation
Sun, J.-M., Wu, T., and Gao, L · 2024
Closest in time.
Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis
Team, K · 2024
Closest in time.
Wu, X., Zhang, D., Gan, R., Lu, J., Wu, Z., Sun, R., Zhang, J., Zhang, P., and Song, Y · 2024
Closest in time.
Dialoguenerf: towards realistic avatar face-to-face conversation video generation
Yan, Y., Zhou, Z., Wang, Z., Gao, J., and Yang, X · 2024
Closest in time.