Fetching the paper…
Reading the bibliography…
Massive web datasets play a key role in the success of large vision-language models like CLIP and Flamingo.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions
D. Zhu, J. Chen, K. Haydarov, X. Shen, W. Zhang, and M. Elhoseiny · 2003
Earlier work this paper cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. Van Der Smagt, D. Cremers, and T. Brox · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Earlier work this paper cites.
Playing for data: Ground truth from computer games
S. R. Richter, V. Vineet, S. Roth, and V. Koltun · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
Visda: The visual domain adaptation challenge
X. Peng, B. Usman, N. Kaushik, J. Hoffman, D. Wang, and K. Saenko · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson · 2019
Earlier work this paper cites.
Fairface: Face attribute dataset for balanced race, gender, and age
K. Kärkkäinen and J. Joo · 2019
Earlier work this paper cites.
Threedworld: A platform for interactive multi-modal physical simulation
C. Gan, J. Schwartz, S. Alter, D. Mrowca, M. Schrimpf, J. Traer, J. De Freitas, J. Kubilius, A. Bhandwaldar, N. Haber, et al · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Earlier work this paper cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
H. Tan and M. Bansal · 2020
Earlier work this paper cites.
Structured3d: A large photo-realistic dataset for structured 3d modeling
J. Zheng, J. Zhang, J. Li, R. Tang, S. Gao, and Z. Zhou · 2020
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
A. Birhane, V. U. Prabhu, and E. Kahembwe · 2021
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Cited alongside, same era.
Next-generation deep learning based on simulators and synthetic data
C. M. de Melo, A. Torralba, L. Guibas, J. DiCarlo, R. Chellappa, and J. Hodgins · 2021
Cited alongside, same era.
Virtex: Learning visual representations from textual annotations
K. Desai and J. Johnson · 2021
Cited alongside, same era.
Redcaps: Web-curated image-text data created by the people, for the people
K. Desai, G. Kaul, Z. Aysola, and J. Johnson · 2021
Cited alongside, same era.
Quality not quantity: On the interaction between dataset design and robustness of clip
T. Nguyen, G. Ilharco, M. Wortsman, S. Oh, and L. Schmidt · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Later among the works it cites.
Is a caption worth a thousand images? a controlled study for representation learning
S. Santurkar, Y. Dubois, R. Taori, P. Liang, and T. Hashimoto · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Cited alongside, same era.
Openclip, July 2021
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Cited alongside, same era.
Combined scaling for zero-shot transfer learning
H. Pham, Z. Dai, G. Ghiasi, K. Kawaguchi, H. Liu, A. W. Yu, J. Yu, Y.-T. Chen, M.-T. Luong, Y. Wu, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Simvlm: Simple visual language model pretraining with weak supervision
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2021
Cited alongside, same era.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Later among the works it cites.
Data feedback loops: Model-driven amplification of dataset biases
R. Taori and T. B. Hashimoto · 2022
Later among the works it cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
A. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos · 2023
Closest in time.
Synthetic data from diffusion models improves imagenet classification
S. Azizi, S. Kornblith, C. Saharia, M. Norouzi, and D. J. Fleet · 2023
Closest in time.
Leaving reality to imagination: Robust classification via generated datasets
H. Bansal and A. Grover · 2023
Closest in time.
Less is more: Removing text-regions improves clip training efficiency and robustness
L. Cao, B. Zhang, C. Chen, Y. Yang, X. Du, W. Zhang, Z. Lu, and Y. Zheng · 2023
Closest in time.
Improving clip training with language rewrites
L. Fan, D. Krishnan, P. Isola, D. Katabi, and Y. Tian · 2023
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
T-mars: Improving visual representations by circumventing text feature learning
P. Maini, S. Goyal, Z. C. Lipton, J. Z. Kolter, and A. Raghunathan · 2023
Closest in time.
Filtering, distillation, and hard negatives for vision-language pre-training
F. Radenovic, A. Dubey, A. Kadian, T. Mihaylov, S. Vandenhende, Y. Patel, Y. Wen, V. Ramanathan, and D. Mahajan · 2023
Closest in time.
Model dementia: Generated data makes models forget
I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot, and R. Anderson · 2023
Closest in time.