Fetching the paper…
Reading the bibliography…
Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
D. Elliott, S. Frank, K. Sima’an, and L. Specia · 2016
Earlier work this paper cites.
Cross-lingual image caption generation
T. Miyazaki and N. Shimizu · 2016
Earlier work this paper cites.
Fluency-guided cross-lingual image captioning
W. Lan, X. Li, and J. Dong · 2017
Earlier work this paper cites.
S. Shankar, Y. Halpern, E. Breck, J. Atwood, J. Wilson, and D. Sculley · 2017
Earlier work this paper cites.
Findings of the third shared task on multimodal machine translation
L. Barrault, F. Bougares, L. Specia, C. Lala, D. Elliott, and S. Frank · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz · 2019
Earlier work this paper cites.
Does object recognition work for everyone?
T. De Vries, I. Misra, C. Wang, and L. Van der Maaten · 2019
Earlier work this paper cites.
Coco-cn for cross-lingual image tagging, captioning, and retrieval
X. Li, C. Xu, X. Wang, W. Lan, Z. Jia, G. Yang, and J. Xu · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. Lipton, and E. P. Xing · 2019
Earlier work this paper cites.
Beyond english-centric multilingual machine translation, 2020
A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, N. Goyal, T. Birch, V. Liptchinsky, S. Edunov, E. Grave, M. Auli, and A. Joulin · 2020
Earlier work this paper cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
N. Reimers and I. Gurevych · 2020
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
xgqa: Cross-lingual visual question answering
J. Pfeiffer, G. Geigle, A. Kamath, J.-M. O. Steitz, S. Roth, I. Vulić, and I. Gurevych · 2021
Cited alongside, same era.
Combined scaling for zero-shot transfer learning, 2021
H. Pham, Z. Dai, G. Ghiasi, H. Liu, A. W. Yu, M.-T. Luong, M. Tan, and Q. V. Le · 2021
Cited alongside, same era.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
K. Pillutla, S. Swayamdipta, R. Zellers, J. Thickstun, S. Welleck, Y. Choi, and Z. Harchaoui · 2021
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
W. A. G. Rojas, S. Diamos, K. R. Kini, D. Kanter, V. J. Reddi, and C. Coleman · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Later among the works it cites.
Chinese clip: Contrastive vision-language pretraining in chinese
A. Yang, J. Pan, J. Lin, R. Men, Y. Zhang, J. Zhou, and C. Zhou · 2022
Later among the works it cites.
A. Fang, A. M. Jose, A. Jain, L. Schmidt, A. Toshev, and V. Shankar · 2023
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Broaden the vision: Geo-diverse visual commonsense reasoning
D. Yin, L. H. Li, Z. Hu, N. Peng, and K.-W. Chang · 2021
Cited alongside, same era.
Winogavil: Gamified association benchmark to challenge vision-and-language models
Y. Bitton, N. Bitton Guetta, R. Yosef, Y. Elovici, M. Bansal, G. Stanovsky, and R. Schwartz · 2022
Cited alongside, same era.
Cross-lingual and multilingual clip
F. Carlsson, P. Eisen, F. Rekathati, and M. Sahlgren · 2022
Cited alongside, same era.
Maxm: Towards multilingual visual question answering
S. Changpinyo, L. Xue, I. Szpektor, A. V. Thapliyal, J. Amelot, M. Yarom, X. Chen, and R. Soricut · 2022
Cited alongside, same era.
Altclip: Altering the language encoder in clip for extended language capabilities
Z. Chen, G. Liu, B.-W. Zhang, F. Ye, Q. Yang, and L. Wu · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, et al · 2022
Cited alongside, same era.
Later among the works it cites.
T-mars: Improving visual representations by circumventing text feature learning
P. Maini, S. Goyal, Z. C. Lipton, J. Z. Kolter, and A. Raghunathan · 2023
Later among the works it cites.
Does progress on object recognition benchmarks improve real-world generalization?
M. Richards, P. Kirichenko, D. Bouchacourt, and M. Ibrahim · 2023
Later among the works it cites.
Nlpositionality: Characterizing design biases of datasets and models
S. Santy, J. T. Liang, R. L. Bras, K. Reinecke, and M. Sap · 2023
Later among the works it cites.
Nllb-clip–train performant multilingual image retrieval model on a budget
A. Visheratin · 2023
Later among the works it cites.
Cultural and linguistic diversity improves visual representations
A. Ye, S. Santy, J. D. Hwang, A. X. Zhang, and R. Krishna · 2023
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2024
Closest in time.
Who’s in and who’s out? a case study of multimodal clip-filtering in datacomp
R. Hong, W. Agnew, T. Kohno, and J. Morgenstern · 2024
Closest in time.
Hype: Hyperbolic entailment filtering for underspecified images and texts
W. Kim, S. Chun, T. Kim, D. Han, and S. Yun · 2024
Closest in time.
Improving multimodal datasets with image captioning
T. Nguyen, S. Y. Gadre, G. Ilharco, S. Oh, and L. Schmidt · 2024
Closest in time.
Geode: a geographically diverse evaluation dataset for object recognition
V. V. Ramaswamy, S. Y. Lin, D. Zhao, A. Adcock, L. van der Maaten, D. Ghadiyaram, and O. Russakovsky · 2024
Closest in time.