Fetching the paper…
Reading the bibliography…
Today, the exponential rise of large models developed by academic and industrial institutions with the help of massive computing resources raises the question of whether someone without access to such resources can make a valuable scientific contribution.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Bridge correlational neural networks for multilingual multimodal representation learning
J. Rajendran, M. M. Khapra, S. Chandar, and B. Ravindran · 2016
Earlier work this paper cites.
On the impact of data set size in transfer learning using deep neural networks
D. Soekhoe, P. Van Der Putten, and A. Plaat · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Coco-cn for cross-lingual image tagging, captioning, and retrieval
X. Li, C. Xu, X. Wang, W. Lan, Z. Jia, G. Yang, and J. Xu · 2019
Earlier work this paper cites.
Impact of training dataset size on neural answer selection models
T. Linjordet and K. Balog · 2019
Earlier work this paper cites.
Towards zero-shot cross-lingual image retrieval, 2020
P. Aggarwal and A. Kale · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
The state and fate of linguistic diversity and inclusion in the nlp world
P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury · 2020
Cited alongside, same era.
Participatory research for low-resourced machine translation: A case study in african languages
W. Nekoto, V. Marivate, T. Matsila, T. Fasubaa, T. Kolawole, T. Fagbohungbe, S. O. Akinola, S. H. Muhammad, S. Kabongo, S. Osei, et al · 2020
Cited alongside, same era.
Contrastive language-image pre-training for the italian language
F. Bianchi, G. Attanasio, R. Pisoni, S. Terragni, G. Sarti, and S. Lakshmi · 2021
Cited alongside, same era.
Mural: Multimodal, multitask representations across languages
A. Jain, M. Guo, K. Srinivasan, T. Chen, S. Kudugunta, C. Jia, Y. Yang, and J. Baldridge · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Clip model is an efficient continual learner
V. Thengane, S. Khan, M. Hayat, and F. Khan · 2022
Later among the works it cites.
Zero and r2d2: A large-scale chinese cross-modal benchmark and a vision-language framework, 2022
C. Xie, H. Cai, J. Song, J. Li, F. Kong, X. Wu, H. Morimitsu, L. Yao, D. Wang, X. Zhang, D. Leng, X. Ji, and Y. Deng · 2022
Later among the works it cites.
Chinese clip: Contrastive vision-language pretraining in chinese
A. Yang, J. Pan, J. Lin, R. Men, Y. Zhang, J. Zhou, and C. Zhou · 2022
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Cross-lingual and multilingual clip
F. Carlsson, P. Eisen, F. Rekathati, and M. Sahlgren · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, et al · 2022
Cited alongside, same era.
8-bit optimizers via block-wise quantization
T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer · 2022
Cited alongside, same era.
Adapting clip for phrase localization without further training
J. Li, G. Shakhnarovich, and R. A. Yeh · 2022
Cited alongside, same era.
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
A. Thapliyal, J. Pont-Tuset, X. Chen, and R. Soricut · 2022
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al
Cited in the paper.
Altclip: Altering the language encoder in clip for extended language capabilities
Z. Chen, G. Liu, B.-W. Zhang, F. Ye, Q. Yang, and L. Wu
Cited in the paper.
Closest in time.
Y. Ji, Y. Deng, Y. Gong, Y. Peng, Q. Niu, L. Zhang, B. Ma, and X. Li · 2023
Closest in time.
Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation
Y. Lin, M. Chen, W. Wang, B. Wu, K. Li, B. Lin, H. Liu, and X. He · 2023
Closest in time.
Laion-coco-nllb dataset, 2023
A. Visheratin · 2023
Closest in time.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Closest in time.
STAIR captions: Constructing a large-scale Japanese image caption dataset
Y. Yoshikawa, Y. Shigeto, and A. Takeuchi · 2066
Closest in time.