Fetching the paper…
Reading the bibliography…
Recent text-to-image generative models, e.g., Stable Diffusion V3 and Flux, have achieved notable progress.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Novel datasets for fine-grained image categorization
Dataset, E · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
A retrieve-and-edit framework for predicting structured outputs
Hashimoto, T. B., Guu, K., Oren, Y., and Liang, P. S · 2018
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Earlier work this paper cites.
Hard negative mixing for contrastive learning
Kalantidis, Y., Sariyildiz, M. B., Pion, N., Weinzaepfel, P., and Larlus, D · 2020
Earlier work this paper cites.
Supervised contrastive learning
Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., and Krishnan, D · 2020
Earlier work this paper cites.
Contrastive representation learning: A framework and review
Le-Khac, P. H., Healy, G., and Smeaton, A. F · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Contrastive learning with hard negative samples
Robinson, J., Chuang, C.-Y., Sra, S., and Jegelka, S · 2020
Earlier work this paper cites.
Clip2video: Mastering video-text retrieval via image clip
Fang, H., Xiong, P., Xu, L., and Chen, Y · 2021
Earlier work this paper cites.
Polyvit: Co-training vision transformers on images, videos and audio
Likhosherstov, V., Arnab, A., Choromanski, K., Lucic, M., Tay, Y., Weller, A., and Dehghani, M · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Cross-modal contrastive learning for text-to-image generation
Zhang, H., Koh, J. Y., Baldridge, J., Lee, H., and Yang, Y · 2021
Earlier work this paper cites.
Retrieval-augmented diffusion models
Blattmann, A., Rombach, R., Oktay, K., Müller, J., and Ommer, B · 2022
Earlier work this paper cites.
Re-imagen: Retrieval-augmented text-to-image generator
Chen, W., Hu, H., Saharia, C., and Cohen, W. W · 2022
Earlier work this paper cites.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2022
Earlier work this paper cites.
Audioclip: Extending clip to image, text and audio
Guzhov, A., Raue, F., Hees, J., and Dengel, A · 2022
Cited alongside, same era.
Clip2point: Transfer clip to point cloud classification with image-depth pre-training
Huang, T., Dong, B., Yang, Y., Huang, X., Lau, R. W., Ouyang, W., and Zuo, W · 2022
Cited alongside, same era.
Liu, Z., Xiong, C., Lv, Y., Liu, Z., and Yu, G · 2022
Cited alongside, same era.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Luo, H., Ji, L., Zhong, M., Chen, Y., Lei, W., Duan, N., and Li, T · 2022
Cited alongside, same era.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Replug: Retrieval-augmented black-box language models
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t · 2023
Later among the works it cites.
Learning audio-visual source localization via false negative aware contrastive learning
Sun, W., Zhang, J., Wang, J., Liu, Z., Zhong, Y., Feng, T., Guo, Y., Zhang, Y., and Barnes, N · 2023
Later among the works it cites.
Any-to-any generation via composable diffusion
Tang, Z., Yang, Z., Zhu, C., Zeng, M., and Bansal, M · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Knn-diffusion: Image generation via large-scale retrieval
Sheynin, S., Ashual, O., Polyak, A., Singer, U., Gafni, O., Nachmani, E., and Taigman, Y · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Cited alongside, same era.
Clip-vip: Adapting pre-trained image-text model to video-language representation alignment
Xue, H., Sun, Y., Liu, B., Fu, J., Song, R., Li, H., and Luo, J · 2022
Cited alongside, same era.
Pointclip: Point cloud understanding by clip
Zhang, R., Guo, Z., Zhang, W., Li, K., Miao, X., Cui, B., Qiao, Y., Gao, P., and Li, H · 2022
Cited alongside, same era.
Componerf: Text-guided multi-object compositional nerf with editable 3d scene layout
Bai, H., Lyu, Y., Jiang, L., Li, S., Lu, H., Lin, X., and Wang, L · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., and Wang, H · 2023
Cited alongside, same era.
Zhu, B., Lin, B., Ning, M., Yan, Y., Cui, J., Wang, H., Pang, Y., Jiang, W., Zhang, J., Li, Z., et al · 2023
Later among the works it cites.
Black forest labs; frontier ai lab, 2024
BlackForest · 2024
Later among the works it cites.
Diffusion forcing: Next-token prediction meets full-sequence diffusion
Chen, B., Monso, D. M., Du, Y., Simchowitz, M., Tedrake, R., and Sitzmann, V · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Later among the works it cites.
Vit-lens: Towards omni-modal representations
Lei, W., Ge, Y., Yi, K., Zhang, J., Gao, D., Sun, D., Ge, Y., Shan, Y., and Shou, M. Z · 2024
Later among the works it cites.
Advances in 3d generation: A survey
Li, X., Zhang, Q., Kang, D., Cheng, W., Gao, Y., Zhang, J., Liang, Z., Liao, J., Cao, Y.-P., and Shan, Y · 2024
Later among the works it cites.
Generative multimodal models are in-context learners
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Wang, Y., Rao, Y., Liu, J., Huang, T., and Wang, X · 2024
Later among the works it cites.
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Tang, J., Chen, Z., Chen, X., Wang, T., Zeng, G., and Liu, Z · 2024
Later among the works it cites.
Omnigen: Unified image generation
Xiao, S., Wang, Y., Zhou, J., Yuan, H., Xing, X., Yan, R., Wang, S., Huang, T., and Liu, Z · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Binding touch to everything: Learning unified multimodal tactile representations
Yang, F., Feng, C., Chen, Z., Park, H., Wang, D., Dou, Y., Zeng, Z., Chen, X., Gangopadhyay, R., Owens, A., et al · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2024
Later among the works it cites.
Diffusion as shader: 3d-aware video diffusion for versatile video generation control
Gu, Z., Yan, R., Lu, J., Li, P., Dou, Z., Si, C., Dong, Z., Liu, Q., Lin, C., Liu, Z., et al · 2025
Closest in time.
Dimer: Disentangled mesh reconstruction model
Jiang, L., Lin, J., Chen, K., Ge, W., Yang, X., Jiang, Y., Lyu, Y., Zheng, X., and Chen, Y · 2025
Closest in time.
From reusing to forecasting: Accelerating diffusion models with taylorseers
Liu, J., Zou, C., Lyu, Y., Chen, J., and Zhang, L · 2025
Closest in time.
Cinemaster: A 3d-aware and controllable framework for cinematic text-to-video generation
Wang, Q., Luo, Y., Shi, X., Jia, X., Lu, H., Xue, T., Wang, X., Wan, P., Zhang, D., and Gai, K · 2025
Closest in time.
Retrieval augmented generation and understanding in vision: A survey and new outlook
Zheng, X., Weng, Z., Lyu, Y., Jiang, L., Xue, H., Ren, B., Paudel, D., Sebe, N., Van Gool, L., and Hu, X · 2025
Closest in time.