Fetching the paper…
Reading the bibliography…
Large-scale cross-modal pre-training paradigms have recently shown ubiquitous success on a wide range of downstream tasks, e.g., zero-shot classification, retrieval and image captioning.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Qi, D.; Su, L.; Song, J.; Cui, E.; Bharti, T.; and Sacheti, A. 2020 · 2001
Earlier work this paper cites.
Nltk: The natural language toolkit
Loper, E.; and Bird, S. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
One-shot learning of object categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2006 · 2006
Earlier work this paper cites.
A study of Gaussian mixture models of color and texture features for image classification and segmentation
Permuter, H.; Francos, J.; Jermyn, I.; et al. 2006 · 2006
Earlier work this paper cites.
Learning from noisy labels with deep neural networks: A survey
Song, H.; Kim, M.; Park, D.; Shin, Y.; and Lee, J.-G. 2020 · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; et al. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J.; Hays, J.; Ehinger, K. A.; Oliva, A.; and Torralba, A. 2010 · 2010
Earlier work this paper cites.
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
Li, W.; Gao, C.; Niu, G.; Xiao, X.; Liu, H.; Liu, J.; Wu, H.; and Wang, H. 2020b · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Gool, L. V. 2014 · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M.; Maji, S.; Kokkinos, I.; Mohamed, S.; and Vedaldi, A. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Reed, S.; Lee, H.; Anguelov, D.; Szegedy, C.; Erhan, D.; and Rabinovich, A. 2014 · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Cited alongside, same era.
YFCC100M: The new data in multimedia research
Thomee, B.; Shamma, D. A.; Friedland, G.; Elizalde, B.; Ni, K.; Poland, D.; Borth, D.; and Li, L.-J. 2016 · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D.; Jastrzębski, S.; Ballas, N.; Krueger, D.; Bengio, E.; Kanwal, M. S.; Maharaj, T.; Fischer, A.; Courville, A.; Bengio, Y.; et al. 2017 · 2017
Cited alongside, same era.
SGDR: Stochastic Gradient Descent with Warm Restarts
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
You, Y.; Li, J.; Reddi, S.; Hseu, J.; Kumar, S.; Bhojanapalli, S.; Song, X.; Demmel, J.; Keutzer, K.; and Hsieh, C.-J. 2020 · 2020
Later among the works it cites.
Error-bounded correction of noisy labels
Zheng, S.; Wu, P.; Goswami, A.; Goswami, M.; Metaxas, D.; and Chen, C. 2020 · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J.; and Gao, J. 2020 · 2020
Later among the works it cites.
Noise estimation using density estimation for self-supervised multimodal learning
Amrani, E.; Ben-Ari, R.; Rotman, D.; and Bronstein, A. 2021 · 2021
Later among the works it cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Loshchilov, I.; and Hutter, F. 2016 · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G.; Tucker, G.; Chorowski, J.; Kaiser, Ł.; and Hinton, G. 2017 · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning
Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
Unsupervised label noise modeling and loss correction
Arazo, E.; Ortego, D.; Albert, P.; O’Connor, N.; and McGuinness, K. 2019 · 2019
Cited alongside, same era.
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2021 · 2021
Later among the works it cites.
Learning with Noisy Correspondence for Cross-modal Matching
Huang, Z.; Niu, G.; Liu, X.; Ding, W.; Xiao, X.; Wu, H.; and Peng, X. 2021 · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q. V.; Sung, Y.; Li, Z.; and Duerig, T. 2021 · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Later among the works it cites.
SLIP: Self-supervision meets Language-Image Pre-training
Mu, N.; Kirillov, A.; Wagner, D.; and Xie, S. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Wang, Z.; Yu, J.; Yu, A. W.; Dai, Z.; Tsvetkov, Y.; and Cao, Y. 2021 · 2021
Later among the works it cites.
FILIP: Fine-grained Interactive Language-Image Pre-Training
Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; and Xu, C. 2021 · 2021
Later among the works it cites.
Florence: A New Foundation Model for Computer Vision
Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. 2021 · 2021
Later among the works it cites.
Cui, Y.; Zhao, L.; Liang, F.; Li, Y.; and Shao, J. 2022 · 2022
Closest in time.
Wukong: 100 Million Large-scale Chinese Cross-modal Pre-training Dataset and A Foundation Framework
Gu, J.; Meng, X.; Lu, G.; Hou, L.; Niu, M.; Xu, H.; Liang, X.; Zhang, W.; Jiang, X.; and Xu, C. 2022 · 2022
Closest in time.
Emphasizing Complementary Samples for Non-Literal Cross-Modal Retrieval
Thomas, C.; and Kovashka, A. 2022 · 2022
Closest in time.
Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
You, H.; Zhou, L.; Xiao, B.; Codella, N.; Cheng, Y.; Xu, R.; Chang, S.-F.; and Yuan, L. 2022 · 2022
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery
Patashnik, O.; Wu, Z.; Shechtman, E.; Cohen-Or, D.; and Lischinski, D. 2021 · 2094
Closest in time.