Fetching the paper…
Reading the bibliography…
Large-scale Vision-Language Models, such as CLIP, learn powerful image-text representations that have found numerous applications, from zero-shot classification to text-to-image generation.
Concerning nonnegative matrices and doubly stochastic matrices
Richard Sinkhorn and Paul Knopp · 1967
Earlier work this paper cites.
Caltech-ucsd birds 200
Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge Belongie, and Pietro Perona · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Comprehension-guided referring expressions
Ruotian Luo and Gregory Shakhnarovich · 2017
Earlier work this paper cites.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Referring expression object segmentation with caption-aware consistency
Yi-Wen Chen, Yi-Hsuan Tsai, Tiantian Wang, Yen-Yu Lin, and Ming-Hsuan Yang · 2019
Earlier work this paper cites.
Learning to assemble neural module tree networks for visual grounding
Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha · 2019
Earlier work this paper cites.
Spair-71k: A large-scale benchmark for semantic correspondence
Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel · 2019
Earlier work this paper cites.
Dynamic graph attention for referring expression comprehension
Sibei Yang, Guanbin Li, and Yizhou Yu · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Learning cross-modal context graph for visual grounding
Yongfei Liu, Bo Wan, Xiaodan Zhu, and Xuming He · 2020
Earlier work this paper cites.
Evaluating clip: towards characterization of broader capabilities and downstream implications
Sandhini Agarwal, Gretchen Krueger, Jack Clark, Alec Radford, Jong Wook Kim, and Miles Brundage · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Zero-shot open set detection by extending clip
Sepideh Esmaeilpour, Bing Liu, Eric Robertson, and Lei Shu · 2021
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Cited alongside, same era.
Referring transformer: A one-step approach to multi-task visual grounding
Muchen Li and Leonid Sigal · 2021
Cited alongside, same era.
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan · 2021
Cited alongside, same era.
Segmentation in style: Unsupervised semantic image segmentation with stylegan and clip
Daniil Pakhomov, Sanchit Hira, Narayani Wagle, Kemar E Green, and Nassir Navab · 2021
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim · 2022
Later among the works it cites.
Pseudo-q: Generating pseudo language queries for visual grounding
Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, and Gao Huang · 2022
Later among the works it cites.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al · 2022
Later among the works it cites.
Open-vocabulary semantic segmentation with mask-adapted clip
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Discriminative triad matching and reconstruction for weakly referring expression grounding
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu, and John Y Goulermas · 2021
Cited alongside, same era.
Are gender-neutral queries really gender-neutral? mitigating gender bias in image search
Jialu Wang, Yang Liu, and Xin Eric Wang · 2021
Cited alongside, same era.
Fake it till you make it: Face analysis in the wild using synthetic data alone, 2021
Erroll Wood, Tadas Baltrušaitis, Charlie Hewitt, Sebastian Dziadzio, Matthew Johnson, Virginia Estellers, Thomas J. Cashman, and Jamie Shotton · 2021
Cited alongside, same era.
Crossing the format boundary of text and boxes: Towards unified vision-language modeling
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Faisal Ahmed, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2021
Cited alongside, same era.
Filip: fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2021
Cited alongside, same era.
Cpt: Colorful prompt tuning for pre-trained vision-language models
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2021
Cited alongside, same era.
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu · 2022
Later among the works it cites.
Disentangling visual and written concepts in clip
Joanna Materzyńska, Antonio Torralba, and David Bau · 2022
Later among the works it cites.
Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization
Luke Melas-Kyriazi, Christian Rupprecht, Iro Laina, and Andrea Vedaldi · 2022
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie · 2022
Later among the works it cites.
Multitask vision-language prompt tuning
Sheng Shen, Shijia Yang, Tianjun Zhang, Bohan Zhai, Joseph E Gonzalez, Kurt Keutzer, and Trevor Darrell · 2022
Later among the works it cites.
Proposalclip: Unsupervised open-category object proposal generation via exploiting clip cues
Hengcan Shi, Munawar Hayat, Yicheng Wu, and Jianfei Cai · 2022
Later among the works it cites.
Reclip: A strong zero-shot baseline for referring expression comprehension
Sanjay Subramanian, Will Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach · 2022
Later among the works it cites.
Cheng-Hao Tu, Zheda Mai, and Wei-Lun Chao · 2022
Later among the works it cites.
Yangtao Wang, Xi Shen, Yuan Yuan, Yuming Du, Maomao Li, Shell Xu Hu, James L Crowley, and Dominique Vaufreydaz · 2022
Later among the works it cites.
Cris: Clip-driven referring image segmentation
Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang · 2022
Later among the works it cites.
Unified vision and language prompt learning
Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy · 2022
Later among the works it cites.
Sheng Zhang, Salman Khan, Zhiqiang Shen, Muzammal Naseer, Guangyi Chen, and Fahad Khan · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Debiasing vision-language models via biased prompts
Ching-Yao Chuang, Varun Jampani, Yuanzhen Li, Antonio Torralba, and Stefanie Jegelka · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.