Fetching the paper…
Reading the bibliography…
Large-scale vision-language pre-training has achieved promising results on downstream tasks.
L. Fei-Fei, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in
2004
Earlier work this paper cites.
L. van der Maaten and G. Hinton, “Viualizing data using t-sne,”
2008
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
A. Krizhevsky, G. Hinton
2009
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “Sun database: Large-scale scene recognition from abbey to zoo,” in
2010
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,”
2011
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in
2012
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,”
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi, “Describing textures in the wild,” in
2014
Earlier work this paper cites.
L. Bossard, M. Guillaumin, and L. Van Gool, “Food-101 – mining discriminative components with random forests,” in
2014
Earlier work this paper cites.
M. Everingham, S. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,”
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,”
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,”
2016
Cited alongside, same era.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in
2017
Cited alongside, same era.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in
2017
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,”
2017
Cited alongside, same era.
A. Van den Oord, Y. Li, O. Vinyals
2018
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in
2018
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,”
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu, “Incorporating convolution designs into visual transformers,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Gildenblat and contributors, “Pytorch library for cam methods,”
2021
Later among the works it cites.
2022
Closest in time.