Fetching the paper…
Reading the bibliography…
The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as GPT-3.
WordNet: An electronic lexical database
George A Miller · 1998
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Cats and dogs
Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi · 2013
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Learning deep representations of fine-grained visual descriptions
Scott Reed, Zeynep Akata, Honglak Lee, and Bernt Schiele · 2016
Earlier work this paper cites.
Link the head to the” beak”: Zero shot learning from noisy text description at part precision
Mohamed Elhoseiny, Yizhe Zhu, Han Zhang, and Ahmed Elgammal · 2017
Earlier work this paper cites.
Fine-grained image classification via combining vision and language
Xiangteng He and Yuxin Peng · 2017
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk · 2018
Earlier work this paper cites.
Towards robust neural machine translation
Yong Cheng, Zhaopeng Tu, Fandong Meng, Junjie Zhai, and Yang Liu · 2018
Earlier work this paper cites.
Hyperspherical variational auto-encoders
Tim R. Davidson, Luca Falorsi, Nicola De Cao, Thomas Kipf, and Jakub M. Tomczak · 2018
Earlier work this paper cites.
Cnn-rnn: a large-scale hierarchical image classification framework
Yanming Guo, Yu Liu, Erwin M Bakker, Yuanhao Guo, and Michael S Lew · 2018
Earlier work this paper cites.
How robust are character-based word embeddings in tagging and mt against wrod scramlbing or randdm nouse?
Georg Heigold, Stalin Varanasi, Günter Neumann, and Josef van Genabith · 2018
Earlier work this paper cites.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Earlier work this paper cites.
Do better imagenet models transfer better?
Simon Kornblith, Jonathon Shlens, and Quoc V. Le · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power
Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing · 2019
Cited alongside, same era.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Mixtext: Linguistically-informed interpolation of hidden space for semi-supervised text classification
Jiaao Chen, Zichao Yang, and Diyi Yang · 2020
Cited alongside, same era.
Evaluating robustness to input perturbations for neural machine translation
What does a platypus look like? generating customized prompts for zero-shot image classification
Sarah Pratt, Rosanne Liu, and Ali Farhadi · 2022
Later among the works it cites.
Integrating language guidance into vision-based deep metric learning
Karsten Roth, Oriol Vinyals, and Zeynep Akata · 2022
Later among the works it cites.
Non-isotropy regularization for proxy-based deep metric learning
Karsten Roth, Oriol Vinyals, and Zeynep Akata · 2022
Later among the works it cites.
To augment or not to augment? a comparative study on text augmentation techniques for low-resource nlp
Gözde Gül Şahin · 2022
Later among the works it cites.
K-lite: Learning transferable visual models with external knowledge
Sheng Shen, Chunyuan Li, Xiaowei Hu, Yujia Xie, Jianwei Yang, Pengchuan Zhang, Anna Rohrbach, Zhe Gan, Lijuan Wang, Lu Yuan, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xing Niu, Prashant Mathur, Georgiana Dinu, and Yaser Al-Onaizan · 2020
Cited alongside, same era.
Zest: Zero-shot learning from text descriptions using textual similarity and visual summarization
Tzuf Paz-Argaman, Reut Tsarfaty, Gal Chechik, and Yuval Atzmon · 2020
Cited alongside, same era.
Mixup-transformer: Dynamic data augmentation for nlp tasks
Lichao Sun, Congying Xia, Wenpeng Yin, Tingting Liang, S Yu Philip, and Lifang He · 2020
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola · 2020
Cited alongside, same era.
Large-scale zero-shot image classification from rich and diverse textual descriptions
Sebastian Bujwid and Josephine Sullivan · 2021
Cited alongside, same era.
The curious layperson: Fine-grained image recognition without expert labels
Subhabrata Choudhury, Iro Laina, Christian Rupprecht, and Andrea Vedaldi · 2021
Cited alongside, same era.
A survey of data augmentation approaches for nlp
Steven Y Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy · 2021
Cited alongside, same era.
Test-time prompt tuning for zero-shot generalization in vision-language models
Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao · 2022
Later among the works it cites.
Sus-x: Training-free name-only transfer of vision-language models
Vishaal Udandarao, Ankush Gupta, and Samuel Albanie · 2022
Later among the works it cites.
Unleashing the power of visual prompting at the pixel level
Junyang Wu, Xianhang Li, Chen Wei, Huiyu Wang, Alan Yuille, Yuyin Zhou, and Cihang Xie · 2022
Later among the works it cites.
Dual modality prompt tuning for vision-language pre-trained model
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang, and Yanning Zhang · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Leaving reality to imagination: Robust classification via generated datasets
Hritik Bansal and Aditya Grover · 2023
Closest in time.
PLOT: Prompt learning with optimal transport for vision-language models
Guangyi Chen, Weiran Yao, Xiangchen Song, Xinyue Li, Yongming Rao, and Kun Zhang · 2023
Closest in time.
Using language to extend to unseen domains, 2023
Lisa Dunlap, Clara Mohri, Devin Guillory, Han Zhang, Trevor Darrell, Joseph E. Gonzalez, Aditi Raghunathan, and Anja Rohrbach · 2023
Closest in time.
Mixgen: A new multi-modal data augmentation
Xiaoshuai Hao, Yi Zhu, Srikar Appalaraju, Aston Zhang, Wanqian Zhang, Bo Li, and Mu Li · 2023
Closest in time.
Is synthetic data from generative models ready for image recognition?
Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue, Wenqing Zhang, Philip Torr, Song Bai, and XIAOJUAN QI · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Closest in time.
Doubly right object recognition: A why prompt for visual rationales, 2023
Chengzhi Mao, Revant Teotia, Amrutha Sundar, Sachit Menon, Junfeng Yang, Xin Wang, and Carl Vondrick · 2023
Closest in time.
Visual classification via description from large language models
Sachit Menon and Carl Vondrick · 2023
Closest in time.
I2mvformer: Large language model generated multi-view document supervision for zero-shot image classification
Muhammad Ferjad Naeem, Muhammad Gul Zain Ali Khan, Yongqin Xian, Muhammad Zeshan Afzal, Didier Stricker, Luc Van Gool, and Federico Tombari · 2023
Closest in time.
Chils: Zero-shot image classification with hierarchical label sets
Zachary Novack, Saurabh Garg, Julian McAuley, and Zachary C Lipton · 2023
Closest in time.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2023
Closest in time.