Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) such as CLIP have shown promising performance on a variety of recognition tasks using the standard zero-shot classification procedure -- computing similarity between the query image and the embedded words for each category.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
On Robustness and Transferability of Convolutional Neural Networks, March 2021
Josip Djolonga, Jessica Yung, Michael Tschannen, Rob Romijnders, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Matthias Minderer, Alexander D’Amour, Dan Moldovan, Sylvain Gelly, Neil Houlsby, Xiaohua Zhai, and Mario Lucic · 2007
Earlier work this paper cites.
In Search of Lost Domain Generalization
Ishaan Gulrajani and David Lopez-Paz · 2007
Earlier work this paper cites.
Robust Pre-Training by Adversarial Contrastive Learning
Ziyu Jiang, Tianlong Chen, Ting Chen, and Zhangyang Wang · 2010
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset, July 2011
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2010
Earlier work this paper cites.
Interactively building a discriminative vocabulary of nameable attributes
Devi Parikh and Kristen Grauman · 2011
Earlier work this paper cites.
Cats and dogs
Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Attribute-Based Classification for Zero-Shot Visual Object Categorization
Christoph H. Lampert, Hannes Nickisch, and Stefan Harmeling · 2013
Earlier work this paper cites.
Zero-Shot Learning Through Cross-Modal Transfer, March 2013
Richard Socher, Milind Ganjoo, Hamsa Sridhar, Osbert Bastani, Christopher D. Manning, and Andrew Y. Ng · 2013
Earlier work this paper cites.
Food-101 – Mining Discriminative Components with Random Forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Describing Textures in the Wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
An Embarrassingly Simple Approach to Zero-Shot Learning
Bernardino Romera-Paredes and Philip H. S. Torr · 2017
Earlier work this paper cites.
Prototypical Networks for Few-shot Learning, June 2017
Jake Snell, Kevin Swersky, and Richard S. Zemel · 2017
Earlier work this paper cites.
Matching Networks for One Shot Learning, December 2017
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2017
Cited alongside, same era.
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2018
Cited alongside, same era.
This Looks Like That: Deep Learning for Interpretable Image Recognition, December 2019
Chaofan Chen, Oscar Li, Chaofan Tao, Alina Jade Barnett, Jonathan Su, and Cynthia Rudin · 2019
Cited alongside, same era.
Explaining Explanations: An Overview of Interpretability of Machine Learning, February 2019
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2019
Cited alongside, same era.
MDETR – Modulated Detection for End-to-End Multi-Modal Understanding, October 2021
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2021
Later among the works it cites.
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA, September 2021
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2019
Cited alongside, same era.
Do Better ImageNet Models Transfer Better?
Simon Kornblith, Jonathon Shlens, and Quoc V. Le · 2019
Cited alongside, same era.
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Cited alongside, same era.
The Inclusive Images Competition
James Atwood, Yoni Halpern, Pallavi Baljekar, Eric Breck, D. Sculley, Pavel Ostyakov, Sergey I. Nikolenko, Igor Ivanov, Roman Solovyev, Weimin Wang, and Miha Skalic · 2020
Cited alongside, same era.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2020
Cited alongside, same era.
Zero-Shot Learning – A Comprehensive Evaluation of the Good, the Bad and the Ugly, September 2020
Yongqin Xian, Christoph H. Lampert, Bernt Schiele, and Zeynep Akata · 2020
Cited alongside, same era.
Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers
Hila Chefer, Shir Gur, and Lior Wolf · 2021
Cited alongside, same era.
Multimodal Neurons in Artificial Neural Networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Cited alongside, same era.
Florence: A New Foundation Model for Computer Vision, November 2021
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang · 2021
Later among the works it cites.
The risks of versatile models: Resolving task ambiguity for vision-language models
Sachit Menon, Ishaan Preetam Chandratreya, and Carl Vondrick · 2022
Closest in time.
Sarah Pratt, Rosanne Liu, and Ali Farhadi · 2022
Closest in time.
NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks, March 2022
Fawaz Sammani, Tanmoy Mukherjee, and Nikos Deligiannis · 2022
Closest in time.
K-LITE: Learning Transferable Visual Models with External Knowledge, April 2022
Sheng Shen, Chunyuan Li, Xiaowei Hu, Yujia Xie, Jianwei Yang, Pengchuan Zhang, Anna Rohrbach, Zhe Gan, Lijuan Wang, Lu Yuan, Ce Liu, Kurt Keutzer, Trevor Darrell, and Jianfeng Gao · 2022
Closest in time.
FLAVA: A Foundational Language And Vision Alignment Model, March 2022
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Closest in time.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt · 2022
Closest in time.
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language, May 2022
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence · 2022
Closest in time.
OPT: Open Pre-trained Transformer Language Models, June 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Closest in time.