Fetching the paper…
Reading the bibliography…
Language-Image Pre-training has demonstrated promising results on zero-shot and few-shot downstream tasks by prompting visual models with natural language prompts.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Rob Fergus Li Fei-Fei and Pietro Perona · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Richard Socher Li-Jia Li Kai Li Jia Deng, Wei Dong and Li Fei-Fei · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Krista A Ehinger-Aude Oliva Jianxiong Xiao, James Hays and Antonio Torralba · 2010
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Amir Roshan Zamir Khurram Soomro and Mubarak Shah · 2012
Earlier work this paper cites.
Cats and dogs
Andrew Zisserman Omkar M Parkhi, Andrea Vedaldi and CV Jawahar · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jia Deng Jonathan Krause, Michael Stark and Li Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Juho Kannala Matthew Blaschko Subhransu Maji, Esa Rahtu and Andrea Vedaldi · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Matthieu Guillaumin Lukas Bossard and Luc Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
Iasonas Kokkinos Sammy Mohamed Mircea Cimpoi, Subhransu Maji and Andrea Vedaldi · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Dollár Tsung-Yi Lin, Michael Maire and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer imageto-sentence models
Chris M Cervantes Juan C Caicedo Julia Hockenmaier Bryan A Plummer, Liwei Wang and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Chris Buehler Damien Teney Mark Johnson Stephen Gould Peter Anderson, Xiaodong He and Lei Zhang · 2018
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Ludwig Schmidt Benjamin Recht, Rebecca Roelofs and Vaishaal Shankar · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
Zachary Lipton Haohan Wang, Songwei Ge and Eric P Xingr · 2019
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Andreas Dengel Patrick Helber, Benjamin Bischke and Damian Borth · 2019
Cited alongside, same era.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan, Eric Wallace, and Sameer Singh · 2020
Cited alongside, same era.
Language models are few-shot learners
Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell Sandhini Agarwal Ariel Herbert-Voss Gretchen Krueger Tom Henighan Rewon Child Aditya Ramesh Daniel M. Ziegler Jeffrey Wu Clemens Winter Christopher Hesse Mark Chen Eric Sigler Mateusz Litwin Scott Gray Benjamin Chess Jack Clark Christopher Berner Sam McCandlish Alec Radford Ilya Sutskever Dario Amodei Tom B. Brown, Benjamin Mann · 2020
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Rami Al-Rfou Brian Lester and Noah Constant · 2021
Cited alongside, same era.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Inductive biases for deep learning of higher-level cognition
Y Bengio A Goyal · 2022
Later among the works it cites.
Coordination among neural modules through a shared global workspace
Alex Lamb Kartikeya Badola Nan Rosemary Ke Nasim Rahaman Jonathan Binas Charles Blundell Michael Mozer Yoshua Bengio Anirudh Goyal, Aniket Didolkar · 2022
Later among the works it cites.
Prompt-aligned gradient for prompt tuning
Yucheng Han Yue Wu Beier Zhu, Yulei Niu and Hanwang Zhang · 2022
Later among the works it cites.
Adaptformer: Adapting vision transformers for scalable visual recognition
Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, and Ping Luo · 2022
Later among the works it cites.
Flamingo: a visual language model for few-shot learning. arxiv preprint
Jean-Baptiste Alayrac et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Norman Mu Saurav Kadavath Frank Wang Evan Dorundo Rahul Desai Tyler Zhu Samyak Parajuli Mike Guo Dawn Song Jacob Steinhardt and Justin Gilmer. Dan Hendrycks, Steven Basart · 2021
Cited alongside, same era.
Natural adversarial examples
Steven Basart Jacob Steinhardt Dan Hendrycks, Kevin Zhao and Dawn Song · 2021
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2021
Cited alongside, same era.
Align and prompt: Video-and-language pre-training with entity prompts
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven C. H. Hoi · 2021
Cited alongside, same era.
Florence: A new foundation model for computer vision
Yi-Ling Chen-Noel Codella Xiyang Dai Jianfeng Gao Houdong Hu Xuedong Huang Boxin Li Chunyuan Li Ce Liu Mengchen Liu Zicheng Liu Yumao Lu Yu Shi Lijuan Wang Jianfeng Wang Bin Xiao Zhen Xiao Jianwei Yang Michael Zeng Luowei Zhou Pengchuan Zhang Lu Yuan, Dongdong Chen · 2021
Cited alongside, same era.
Improving coherence and consistency in neural sequence models with dual-system, neuro-symbolic reasoning
Joshua B. Tenenbaum Brenden M. Lake Maxwell Nye, Michael Henry Tessler · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
How much can clip benefit vision-and-language tasks?
Hao Tan Mohit Bansal Anna Rohrbach Kai-Wei Chang Zhewei Yao Sheng Shen, Liunian Harold Li and Kurt Keutzer · 2021
Cited alongside, same era.
Language models show human-like content effects on reasoning
Stephanie C. Y. Chan Antonia Creswell Dharshan Kumaran James L. McClelland Felix Hill Ishita Dasgupta, Andrew K. Lampinen · 2022
Later among the works it cites.
Prompt distribution learning
Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian · 2022
Later among the works it cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Tony Xia Liang Qiu Kai-Wei Chang Song-Chun Zhu Oyvind Tafjord Peter Clark Ashwin Kalyan Pan Lu, Swaroop Mishra · 2022
Later among the works it cites.
Learning to prompt for continual learning
Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
Ted Xiao Harris Chan Jacky Liang Pete Florence Andy Zeng Jonathan Tompson Igor Mordatch Yevgen Chebotar Pierre Sermanet Noah Brown Tomas Jackson Linda Luu Sergey Levine Karol Hausman Brian Ichter Wenlong Huang, Fei Xia · 2022
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.