Fetching the paper…
Reading the bibliography…
Understanding long text is of great demands in practice but beyond the reach of most language-image pre-training (LIP) models.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Learning to assemble neural module tree networks for visual grounding
D. Liu, H. Zhang, F. Wu, and Z.-J. Zha · 2019
Earlier work this paper cites.
Context-aware visual policy network for fine-grained image captioning
Z.-J. Zha, D. Liu, H. Zhang, Y. Zhang, and F. Wu · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Learning to discretely compose reasoning module networks for video captioning
G. Tan, D. Liu, M. Wang, and Z.-J. Zha · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Cited alongside, same era.
Structured multi-level interaction network for video moment localization via language query
H. Wang, Z.-J. Zha, L. Li, D. Liu, and J. Luo · 2021
Cited alongside, same era.
FILIP: Fine-grained interactive language-image pre-training
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu · 2021
Cited alongside, same era.
Coyo-700m: Image-text pair dataset
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim · 2022
Cited alongside, same era.
Scaling language-image pre-training via masking
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He · 2023
Later among the works it cites.
StableRep: Synthetic images from text-to-image models make strong visual representation learners
Y. Tian, L. Fan, P. Isola, H. Chang, and D. Krishnan · 2023
Later among the works it cites.
A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions
J. Urbanek, F. Bordes, P. Astolfi, M. Williamson, V. Sharma, and A. Romero-Soriano · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
W. Xiong, J. Liu, I. Molybog, H. Zhang, P. Bhargava, R. Hou, L. Martin, R. Rungta, K. A. Sankararaman, B. Oguz, et al · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
UniCLIP: Unified framework for contrastive language-image pre-training
J. Lee, J. Kim, H. Shon, B. Kim, S. H. Kim, H. Lee, and J. Kim · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
LiT: Zero-shot transfer with locked-image text tuning
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer · 2022
Cited alongside, same era.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi · 2023
Cited alongside, same era.
Improving CLIP training with language rewrites
L. Fan, D. Krishnan, P. Isola, D. Katabi, and Y. Tian · 2023
Cited alongside, same era.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Cited alongside, same era.
RLEG: vision-language representation learning with diffusion-based embedding generation
L. Zhao, K. Zheng, Y. Zheng, D. Zhao, and J. Zhou · 2023
Later among the works it cites.
Vision transformers need registers
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski · 2024
Closest in time.
Imageinwords: Unlocking hyper-detailed image descriptions
R. Garg, A. Burns, B. K. Ayan, Y. Bitton, C. Montgomery, Y. Onoe, A. Bunner, R. Krishna, J. Baldridge, and R. Soricut · 2024
Closest in time.
SynthCLIP: Are we ready for a fully synthetic clip training?
H. A. A. K. Hammoud, H. Itani, F. Pizzati, P. Torr, A. Bibi, and B. Ghanem · 2024
Closest in time.
Long-clip: Unlocking the long-text capability of clip
B. Zhang, P. Zhang, X. Dong, Y. Zang, and J. Wang · 2024
Closest in time.
DreamLIP: Language-image pre-training with long captions
K. Zheng, Y. Zhang, W. Wu, F. Lu, S. Ma, X. Jin, W. Chen, and Y. Shen · 2024
Closest in time.