Fetching the paper…
Reading the bibliography…
Text-to-image diffusion models can synthesize high-quality images, but they have various limitations.
Who belongs in the family?
Thorndike, R. L. 1953 · 1953
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Efron, B. 1992 · 1992
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.; Kai Li; and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
Wah, C.; Branson, S.; Welinder, P.; Perona, P.; and Belongie, S. 2011 · 2011
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Matching Networks for One Shot Learning
Vinyals, O.; Blundell, C.; Lillicrap, T.; kavukcuoglu, k.; and Wierstra, D. 2016 · 2016
Earlier work this paper cites.
Apache spark: a unified engine for big data processing
Zaharia, M.; Xin, R. S.; Wendell, P.; Das, T.; Armbrust, M.; Dave, A.; Meng, X.; Rosen, J.; Venkataraman, S.; Franklin, M. J.; et al. 2016 · 2016
Earlier work this paper cites.
Multi-Level Semantic Feature Augmentation for One-Shot Learning
Chen, Z.; Fu, Y.; Zhang, Y.; Jiang, Y.-G.; Xue, X.; and Sigal, L. 2018 · 2018
Earlier work this paper cites.
On gans and gmms
Richardson, E.; and Weiss, Y. 2018 · 2018
Earlier work this paper cites.
The inaturalist species classification and detection dataset
Van Horn, G.; Mac Aodha, O.; Song, Y.; Cui, Y.; Sun, C.; Shepard, A.; Adam, H.; Perona, P.; and Belongie, S. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018 · 2018
Earlier work this paper cites.
Meta-learning with differentiable closed-form solvers
Bertinetto, L.; Henriques, J. F.; Torr, P. H.; and Vedaldi, A. 2019 · 2019
Earlier work this paper cites.
Large-scale long-tailed recognition in an open world
Liu, Z.; Miao, Z.; Zhan, X.; Wang, J.; Gong, B.; and Yu, S. X. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Decoupling Representation and Classifier for Long-Tailed Recognition
Kang, B.; Xie, S.; Rohrbach, M.; Yan, M.; Gordo, A.; Feng, J.; and Kalantidis, Y. 2020 · 2020
Earlier work this paper cites.
Supervised contrastive learning
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020 · 2020
Earlier work this paper cites.
Reliable Fidelity and Diversity Metrics for Generative Models
Naeem, M. F.; Oh, S. J.; Uh, Y.; Choi, Y.; and Yoo, J. 2020 · 2020
Earlier work this paper cites.
Few-Shot Learning via Embedding Adaptation With Set-to-Set Functions
Ye, H.-J.; Hu, H.; Zhan, D.-C.; and Sha, F. 2020 · 2020
Earlier work this paper cites.
DeepEMD: Few-Shot Image Classification With Differentiable Earth Mover’s Distance and Structured Classifiers
Zhang, C.; Cai, Y.; Lin, G.; and Shen, C. 2020 · 2020
Earlier work this paper cites.
Parametric Contrastive Learning
Cui, J.; Zhong, Z.; Liu, S.; Yu, B.; and Jia, J. 2021 · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2021 · 2021
Cited alongside, same era.
MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition
Li, S.; Gong, K.; Liu, C. H.; Wang, Y.; Qiao, F.; and Cheng, X. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
From Generalized Zero-Shot Learning to Long-Tail With Class Descriptors
Samuel, D.; Atzmon, Y.; and Chechik, G. 2021 · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.; Ghasemipour, S. K. S.; Ayan, B. K.; Mahdavi, S. S.; Lopes, R. G.; et al. 2022 · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Later among the works it cites.
Baby steps towards few-shot learning with multiple semantics
Schwartz, E.; Karlinsky, L.; Feris, R.; Giryes, R.; and Bronstein, A. 2022 · 2022
Later among the works it cites.
MaxViT: Multi-Axis Vision Transformer
Tu, Z.; Talebi, H.; Zhang, H.; Yang, F.; Milanfar, P.; Bovik, A.; and Li, Y. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Distributional Robustness Loss for Long-tail Learning
Samuel, D.; and Chechik, G. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition
Tian, C.; Wang, W.; Zhu, X.; Wang, X.; Dai, J.; and Qiao, Y. 2021 · 2021
Cited alongside, same era.
Long-tailed Recognition by Routing Diverse Distribution-Aware Experts
Wang, X.; Lian, L.; Miao, Z.; Liu, Z.; and Yu, S. X. 2021 · 2021
Cited alongside, same era.
Deep long-tailed learning: A survey
Zhang, Y.; Kang, B.; Hooi, B.; Yan, S.; and Feng, J. 2021 · 2021
Cited alongside, same era.
Learning to Prompt for Vision-Language Models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2021 · 2021
Cited alongside, same era.
Wang, Z. J.; Montoya, E.; Munechika, D.; Yang, H.; Hoover, B.; and Chau, D. H. 2022 · 2022
Later among the works it cites.
Generating Representative Samples for Few-Shot Classification
Xu, J.; and Le, H. 2022 · 2022
Later among the works it cites.
SEGA: semantic guided attention on visual prototype for few-shot learning
Yang, F.; Wang, R.; and Chen, X. 2022 · 2022
Later among the works it cites.
Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
Zhang, R.; Fang, R.; Zhang, W.; Gao, P.; Li, K.; Dai, J.; Qiao, Y. J.; and Li, H. 2022 · 2022
Later among the works it cites.
SpaText: Spatio-Textual Representation for Controllable Image Generation
Avrahami, O.; Hayes, T.; Gafni, O.; Gupta, S.; Taigman, Y.; Parikh, D.; Lischinski, D.; Fried, O.; and Yin, X. 2023 · 2023
Closest in time.
Synthetic Data from Diffusion Models Improves ImageNet Classification
Azizi, S.; Kornblith, S.; Saharia, C.; Norouzi, M.; and Fleet, D. J. 2023 · 2023
Closest in time.
Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models
Chefer, H.; Alaluf, Y.; Vinker, Y.; Wolf, L.; and Cohen-Or, D. 2023 · 2023
Closest in time.
Fine-grained Visual Classification with High-temperature Refinement and Background Suppression
Chou, P.-Y.; Kao, Y.-Y.; and Lin, C.-H. 2023 · 2023
Closest in time.
Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Feng, W.; He, X.; Fu, T.-J.; Jampani, V.; Akula, A.; Narayana, P.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2023 · 2023
Closest in time.
Is synthetic data from generative models ready for image recognition?
He, R.; Sun, S.; Yu, X.; Xue, C.; Zhang, W.; Torr, P. H. S.; Bai, S.; and Qi, X. 2023 · 2023
Closest in time.
Lian, L.; Li, B.; Yala, A.; and Darrell, T. 2023 · 2023
Closest in time.
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
Rassin, R.; Hirsch, E.; Glickman, D.; Ravfogel, S.; Goldberg, Y.; and Chechik, G. 2023 · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Closest in time.
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
Ryali, C. K.; Hu, Y.-T.; Bolya, D.; Wei, C.; Fan, H.; Huang, P.-Y. B.; Aggarwal, V.; Chowdhury, A.; Poursaeed, O.; Hoffman, J.; Malik, J.; Li, Y.; and Feichtenhofer, C. 2023 · 2023
Closest in time.
Key-Locked Rank One Editing for Text-to-Image Personalization
Tewel, Y.; Gal, R.; Chechik, G.; and Atzmon, Y. 2023 · 2023
Closest in time.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L.; and Agrawala, M. 2023 · 2023
Closest in time.