Fetching the paper…
Reading the bibliography…
Recently, the development of pre-trained vision language foundation models (VLFMs) has led to remarkable performance in many tasks.
Bleu: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Image change detection algorithms: a systematic survey
Radke, R. J.; Andra, S.; Al-Kofahi, O.; and Roysam, B. 2005 · 2005
Earlier work this paper cites.
A large-scale benchmark dataset for event recognition in surveillance video
Oh, S.; Hoogs, A.; Perera, A.; Cuntoor, N.; Chen, C.-C.; Lee, J. T.; Mukherjee, S.; Aggarwal, J.; Lee, H.; Davis, L.; et al. 2011 · 2011
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
Dosovitskiy, A.; Fischer, P.; Ilg, E.; Hausser, P.; Hazirbas, C.; Golkov, V.; Van Der Smagt, P.; Cremers, D.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Spatial transformer networks
Jaderberg, M.; Simonyan, K.; Zisserman, A.; et al. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016 · 2016
Earlier work this paper cites.
Image difference captioning with instance-level fine-grained feature representation
Huang, Q.; Liang, Y.; Wei, J.; Cai, Y.; Liang, H.; Leung, H.-f.; and Li, Q. 2021 · 2017
Earlier work this paper cites.
Flownet 2.0: Evolution of optical flow estimation with deep networks
Ilg, E.; Mayer, N.; Saikia, T.; Keuper, M.; Dosovitskiy, A.; and Brox, T. 2017 · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J.; Hariharan, B.; Van Der Maaten, L.; Fei-Fei, L.; Lawrence Zitnick, C.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Image caption with global-local attention
Li, L.; Tang, S.; Deng, L.; Zhang, Y.; and Tian, Q. 2017 · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Deep feature flow for video recognition
Zhu, X.; Xiong, Y.; Dai, J.; Yuan, L.; and Wei, Y. 2017 · 2017
Cited alongside, same era.
Learning to describe differences between pairs of similar images
Jhamtani, H.; and Berg-Kirkpatrick, T. 2018 · 2018
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Cited alongside, same era.
Robust change captioning
Park, D. H.; Darrell, T.; and Rohrbach, A. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Simmim: A simple framework for masked image modeling
Xie, Z.; Zhang, Z.; Cao, Y.; Lin, Y.; Bao, J.; Yao, Z.; Dai, Q.; and Hu, H. 2022 · 2022
Later among the works it cites.
Image difference captioning with pre-training and contrastive learning
Yao, L.; Wang, W.; and Jin, Q. 2022 · 2022
Later among the works it cites.
Changes to Captions: An Attentive Network for Remote Sensing Change Captioning
Chang, S.; and Ghamisi, P. 2023 · 2023
Closest in time.
SARAS-net: scale and relation aware siamese network for change detection
Chen, C.-P.; Hsieh, J.-W.; Chen, P.-Y.; Hsieh, Y.-K.; and Wang, B.-S. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Semantic flow for fast and accurate scene parsing
Li, X.; You, A.; Zhu, Z.; Zhao, H.; Yang, M.; Yang, K.; Tan, S.; and Tong, Y. 2020 · 2020
Cited alongside, same era.
Finding it at another side: A viewpoint-adapted matching encoder for change captioning
Shi, X.; Yang, X.; Gu, J.; Joty, S.; and Cai, J. 2020 · 2020
Cited alongside, same era.
Image change captioning by learning from an auxiliary task
Hosseinzadeh, M.; and Wang, Y. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Closest in time.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Closest in time.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Neighborhood Contrastive Transformer for Change Captioning
Tu, Y.; Li, L.; Su, L.; Lu, K.; and Huang, Q. 2023 · 2023
Closest in time.
Aim: Adapting image models for efficient video action recognition
Yang, T.; Zhu, Y.; Xie, Y.; Zhang, A.; Chen, C.; and Li, M. 2023 · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023 · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2057
Closest in time.
Agnostic change captioning with cycle consistency
Kim, H.; Kim, J.; Lee, H.; Park, H.; and Kim, G. 2021 · 2095
Closest in time.