Fetching the paper…
Reading the bibliography…
As text-to-image (T2I) synthesis models increase in size, they demand higher inference costs due to the need for more expensive GPUs with larger memory, which makes it challenging to reproduce these models in addition to the restricted access to training datasets.
Matching local self-similarities across images and videos
Shechtman, E. and Irani, M · 2007
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Principal component analysis: a review and recent developments
Jolliffe, I. T. and Cadima, J · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
A comprehensive overhaul of feature distillation
Heo, B., Kim, J., Yun, S., Park, H., Kwak, N., and Choi, J. Y · 2019
Earlier work this paper cites.
Style transfer by relaxed optimal transport and self-similarity
Kolkin, N., Salavon, J., and Shakhnarovich, G · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y · 2021
Earlier work this paper cites.
Openclip
Ilharco, G., Wortsman, M., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
A study on the evaluation of generative models
Betzalel, E., Penso, C., Navon, A., and Fetaya, E · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
Salimans, T. and Ho, J · 2022
Cited alongside, same era.
Splicing vit features for semantic appearance transfer
Tumanyan, N., Bar-Tal, O., Bagon, S., and Dekel, T · 2022
Cited alongside, same era.
Diffusers: State-of-the-art diffusion models
von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., and Wolf, T · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Cited alongside, same era.
Improving image generation with better captions
Laion pop: 600,000 high-resolution images with detailed descriptions
Schuhmann, C. and Bevan, P · 2023
Closest in time.
Post-training quantization on diffusion models
Shang, Y., Yuan, Z., Xie, B., Wu, B., and Yan, Y · 2023
Closest in time.
Mvdream: Multi-view diffusion for 3d generation
Shi, Y., Wang, P., Ye, J., Long, M., Li, K., and Yang, X · 2023
Closest in time.
Plug-and-play diffusion features for text-driven image-to-image translation
Tumanyan, N., Geyer, M., Bagon, S., and Dekel, T · 2023
Closest in time.
Diffusers: State-of-the-art diffusion models
von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., and Wolf, T · 2023
Closest in time.
Cogvlm: Visual expert for pretrained language models, 2023
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., Xu, J., Xu, B., Li, J., Dong, Y., Ding, M., and Tang, J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., Manassra, W., Dhariwal, P., Chu, C., Jiao, Y., and Ramesh, A · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Cited alongside, same era.
Sdxl-vae-fp16-fix
Bohan, O. B · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X · 2023
Cited alongside, same era.
On architectural compression of text-to-image diffusion models
Kim, B.-K., Song, H.-K., Castells, T., and Choi, S · 2023
Cited alongside, same era.
Came: Confidence-guided adaptive memory efficient optimization
Luo, Y., Ren, X., Zheng, Z., Jiang, Z., Jiang, X., and You, Y · 2023
Cited alongside, same era.
On distillation of guided diffusion models
Meng, C., Rombach, R., Gao, R., Kingma, D., Ermon, S., Ho, J., and Salimans, T · 2023
Cited alongside, same era.
Closest in time.
Wu, X., Hao, Y., Sun, K., Chen, Y., Zhu, F., Zhao, R., and Li, H · 2023
Closest in time.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
Zhang, S., Wang, J., Zhang, Y., Zhao, K., Yuan, H., Qin, Z., Wang, X., Zhao, D., and Zhou, J · 2023
Closest in time.
The unreasonable ineffectiveness of the deeper layers
Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., and Roberts, D. A · 2024
Closest in time.
Progressive knowledge distillation of stable diffusion xl using layer level loss, 2024
Gupta, Y., Jaddipal, V. V., Prabhala, H., Paul, S., and Platen, P. V · 2024
Closest in time.
Sdxl-lightning: Progressive adversarial diffusion distillation
Lin, S., Wang, A., and Yang, X · 2024
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect
Men, X., Xu, M., Zhang, Q., Wang, B., Lin, H., Lu, Y., Han, X., and Chen, W · 2024
Closest in time.
Phased consistency model
Wang, F.-Y., Huang, Z., Bergman, A. W., Shen, D., Gao, P., Lingelbach, M., Sun, K., Bian, W., Song, G., Liu, Y., Li, H., and Wang, X · 2024
Closest in time.
Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms
Yang, L., Yu, Z., Meng, C., Xu, M., Ermon, S., and Cui, B · 2024
Closest in time.