Fetching the paper…
Reading the bibliography…
Large-scale text-to-image diffusion models have shown impressive capabilities for generative tasks by leveraging strong vision-language alignment from pre-training.
“Deep unsupervised learning using nonequilibrium thermodynamics,”
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli, · 2015
Earlier work this paper cites.
“U-Net: Convolutional networks for biomedical image segmentation,”
Olaf Ronneberger, Philipp Fischer, and Thomas Brox, · 2015
Earlier work this paper cites.
“Faster R-CNN: Towards real-time object detection with region proposal networks,”
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun, · 2015
Earlier work this paper cites.
“An improved non-monotonic transition system for dependency parsing,”
Matthew Honnibal and Mark Johnson, · 2015
Earlier work this paper cites.
“Modeling context in referring expressions,”
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg, · 2016
Earlier work this paper cites.
“Generation and comprehension of unambiguous object descriptions,”
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy, · 2016
Earlier work this paper cites.
“MAttNet: Modular attention network for referring expression comprehension,”
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg, · 2018
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Earlier work this paper cites.
“VinVL: Revisiting visual representations in vision-language models,”
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao, · 2021
Cited alongside, same era.
“CPT: Colorful prompt tuning for pre-trained vision-language models,”
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun, · 2021
Cited alongside, same era.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Cited alongside, same era.
“LAION-5B: An open large-scale dataset for training next generation image-text models,”
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al., · 2022
Cited alongside, same era.
“GPT-4 technical report,” 2023
OpenAI, · 2023
Cited alongside, same era.
“Prompt-based context-and domain-aware pretraining for vision and language navigation,”
Ting Liu, Wansen Wu, Yue Hu, Youkai Wang, Kai Xu, and Quanjun Yin, · 2023
Closest in time.
“Language adaptive weight generation for multi-task visual grounding,”
Wei Su, Peihan Miao, Huanzhang Dou, Gaoang Wang, Liang Qiao, Zheyang Li, and Xi Li, · 2023
Closest in time.
“DQ-DETR: Dual query detection transformer for phrase extraction and grounding,”
Shilong Liu, Shijia Huang, Feng Li, Hao Zhang, Yaoyuan Liang, Hang Su, Jun Zhu, and Lei Zhang, · 2023
Closest in time.
“Zero-shot referring image segmentation with global-local context features,”
Seonghoon Yu, Paul Hongsuck Seo, and Jeany Son, · 2023
Closest in time.
“Your diffusion model is secretly a zero-shot classifier,”
Alexander C Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Imagic: Text-based real image editing with diffusion models,”
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani, · 2023
Cited alongside, same era.
“SmartBrush: Text and shape guided object inpainting with diffusion model,”
Shaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz, and Kun Zhang, · 2023
Cited alongside, same era.
“Dap: Domain-aware prompt learning for vision-and-language navigation,”
Ting Liu, Yue Hu, Wansen Wu, Youkai Wang, Kai Xu, and Quanjun Yin, · 2023
Cited alongside, same era.
“Unleashing text-to-image diffusion models for visual perception,”
Wenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu, Jie Zhou, and Jiwen Lu, · 2023
Closest in time.
“Diffusion models for zero-shot open-vocabulary segmentation,”
Laurynas Karazija, Iro Laina, Andrea Vedaldi, and Christian Rupprecht, · 2023
Closest in time.
“Discriminative diffusion models as few-shot vision and language learners,”
Xuehai He, Weixi Feng, Tsu-Jui Fu, Varun Jampani, Arjun Akula, Pradyumna Narayana, Sugato Basu, William Yang Wang, and Xin Eric Wang, · 2023
Closest in time.