Fetching the paper…
Reading the bibliography…
Using natural language as a supervision for training visual recognition models holds great promise.
Nltk: The natural language toolkit
Edward Loper and Steven Bird · 2002
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc' Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Richard Socher, Milind Ganjoo, Christopher D Manning, and Andrew Ng · 2013
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens Van Der Maaten, Allan Jabri, and Nicolas Vasilache · 2016
Earlier work this paper cites.
Learning visual n-grams from web data
Ang Li, Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach
Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, and In So Kweon · 2019
Cited alongside, same era.
Data efficient masked language modeling for vision and language
Yonatan Bitton, Gabriel Stanovsky, Michael Elhadad, and Roy Schwartz · 2021
Closest in time.
Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and-language berts
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott · 2021
Closest in time.
Data-efficient language-supervised zero-shot learning with self-distillation
Ruizhe Cheng, Bichen Wu, Peizhao Zhang, Peter Vajda, and Joseph E Gonzalez · 2021
Closest in time.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Closest in time.
With a little help from my friends: Nearest-neighbor contrastive learning of visual representations
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H. Lampert, Bernt Schiele, and Zeynep Akata · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Bootstrap your own latent - a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, koray kavukcuoglu, Remi Munos, and Michal Valko · 2020
Cited alongside, same era.
Stella Frank, Emanuele Bugliarello, and Desmond Elliott · 2021
Closest in time.
DeCLUTR: Deep contrastive learning for unsupervised textual representations
John Giorgi, Osvald Nitski, Bo Wang, and Gary Bader · 2021
Closest in time.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Closest in time.
Mean shift for self-supervised learning
Soroush Abbasi Koohpayegani, Ajinkya Tejankar, and Hamed Pirsiavash · 2021
Closest in time.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi · 2021
Closest in time.
Unsupervised vision-and-language pre-training without parallel images and captions
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, and Kai-Wei Chang · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela · 2021
Closest in time.
UnNatural Language Inference
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams · 2021
Closest in time.
Jue Wang, Haofan Wang, Jincan Deng, Weijia Wu, and Debing Zhang · 2021
Closest in time.
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Anonymous · 2022
Closest in time.