Fetching the paper…
Reading the bibliography…
Recent advances in text-to-image (T2I) diffusion models have enabled impressive image generation capabilities guided by text prompts.
Canny, J.: A computational approach to edge detection. IEEE Transactions on Pattern Analysis and Machine Intelligence p. 679–698 (Jan 2009). https://doi.org/10.1109/tpami.1986.4767851, http://dx.doi.org/10.1109/tpami.1986.4767851
1986
Earlier work this paper cites.
Kingma, D., Welling, M.: Auto-encoding variational bayes. arXiv: Machine Learning,arXiv: Machine Learning (Dec 2013)
2013
Earlier work this paper cites.
Xie, S., Tu, Z.: Holistically-nested edge detection. In: 2015 IEEE International Conference on Computer Vision (ICCV) (Feb 2016). https://doi.org/10.1109/iccv.2015.164, http://dx.doi.org/10.1109/iccv.2015.164
2015
Earlier work this paper cites.
Pont-Tuset, J., Perazzi, F., Caelles, S., Arbeláez, P., Sorkine-Hornung, A., Gool, L.: The 2017 davis challenge on video object segmentation (Apr 2017)
2017
Earlier work this paper cites.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial networks. Communications of the ACM 63
2020
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. (Jan 2020)
2020
Earlier work this paper cites.
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., Koltun, V.: Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence p. 1623–1637 (Aug 2020). https://doi.org/10.1109/tpami.2020.3019967, http://dx.doi.org/10.1109/tpami.2020.3019967
2020
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models (Oct 2020)
2020
Earlier work this paper cites.
Teed, Z., Deng, J.: RAFT: Recurrent All-Pairs Field Transforms for Optical Flow, p. 402–419 (Jan 2020). https://doi.org/10.1007/978-3-030-58536-5_24, http://dx.doi.org/10.1007/978-3-030-58536-5_24
2020
Earlier work this paper cites.
Teed, Z., Deng, J.: RAFT: recurrent all-pairs field transforms for optical flow. In: ECCV (2). Lecture Notes in Computer Science, vol. 12347, pp. 402–419. Springer (2020)
2020
Earlier work this paper cites.
Bain, M., Nagrani, A., Varol, G., Zisserman, A.: Frozen in time: A joint video and image encoder for end-to-end retrieval. In: IEEE International Conference on Computer Vision (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision (Feb 2021)
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
He, Y., Yang, T., Zhang, Y., Shan, Y., Chen, Q.: Latent video diffusion models for high-fidelity video generation with arbitrary lengths (Nov 2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Geyer, M., Bar-Tal, O., Bagon, S., Dekel, T.: Tokenflow: Consistent diffusion features for consistent video editing (Jul 2023)
2023
Closest in time.
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., Dai, B.: Animatediff: Animate your personalized text-to-image diffusion models without specific tuning (Jul 2023)
2023
Closest in time.
Huang, L., Chen, D., Liu, Y., Shen, Y., Zhao, D., Zhou, J.: Composer: Creative and controllable image synthesis with composable conditions (Feb 2023)
2023
Closest in time.
2023
Closest in time.
Liu, S., Zhang, Y., Li, W., Lin, Z., Jia, J.: Video-p2p: Video editing with cross-attention control (Mar 2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Sep 2022). https://doi.org/10.1109/cvpr52688.2022.01042, http://dx.doi.org/10.1109/cvpr52688.2022.01042
2022
Cited alongside, same era.
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., Aberman, K.: Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation (Aug 2022)
2022
Cited alongside, same era.
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al.: Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., Jitsev, J.: Laion-5b: An open large-scale dataset for training next generation image-text models (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Wang, Xiang, e.a.: Videocomposer: Compositional video synthesis with motion controllability. ICCV, 2023
2023
Closest in time.
Wang, X., Yuan, H., Zhang, S., Chen, D., Wang, J., Zhang, Y., Shen, Y., Zhao, D., Zhou, J.: Videocomposer: Compositional video synthesis with motion controllability (Jun 2023)
2023
Closest in time.
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., Dong, Y.: Imagereward: Learning and evaluating human preferences for text-to-image generation (Apr 2023)
2023
Closest in time.
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., Dong, Y.: Imagereward: Learning and evaluating human preferences for text-to-image generation (Apr 2023)
2023
Closest in time.