Fetching the paper…
Reading the bibliography…
Trained on web-scale image-text pairs, Vision-Language Models (VLMs) such as CLIP can recognize images of common objects in a zero-shot fashion.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman. 2008 · 2008
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. 2011 · 2011
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. 2013 · 2013
Earlier work this paper cites.
Link the head to the "beak": Zero shot learning from noisy text description at part precision
Mohamed Elhoseiny, Yizhe Zhu, Han Zhang, and Ahmed Elgammal. 2017 · 2017
Earlier work this paper cites.
Learning visual n-grams from web data
A. Li, A. Jabri, A. Joulin, and L. van der Maaten. 2017 · 2017
Earlier work this paper cites.
Zero-shot learning — a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. 2018 · 2018
Earlier work this paper cites.
Openclip
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, et al. 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, et al. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Large-scale zero-shot learning in the wild: Classifying zoological illustrations
Lise Stork, Andreas Weber, Jaap van den Herik, Aske Plaat, Fons Verbeek, and Katherine Wolstencroft. 2021 · 2021
Cited alongside, same era.
A realistic evaluation of semi-supervised learning for fine-grained classification
Jong-Chyi Su, Zezhou Cheng, and Subhransu Maji. 2021 · 2021
Visual classification via description from large language models
Sachit Menon and Carl Vondrick. 2022 · 2022
Later among the works it cites.
Zero-shot bird species recognition by learning from field guides
Andr’es C. Rodr’iguez, Stefano D’aronco, Rodrigo Caye Daudt, Jan Dirk Wegner, and Konrad Schindler. 2022 · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, et al. 2023 · 2023
Closest in time.
What does a platypus look like? generating customized prompts for zero-shot image classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language-aware soft prompting for vision & language foundation models
Adrian Bulat and Georgios Tzimiropoulos. 2022 · 2022
Cited alongside, same era.
The semi-supervised inaturalist-aves challenge at FGVC7 workshop
Jong-Chyi Su and Subhransu Maji. 2021a
Cited in the paper.
The semi-supervised inaturalist challenge at the FGVC8 workshop
Jong-Chyi Su and Subhransu Maji. 2021b
Cited in the paper.
Sarah Pratt, Ian Covert, Rosanne Liu, and Ali Farhadi. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, et al. 2023 · 2023
Closest in time.